← Marketing issues

Common Crawl opened 580,000 llms.txt files. Two-thirds were empty templates.

Common Crawl, the nonprofit that publishes web-crawl data, analyzed 584,107 llms.txt files found in the wild and published the results on September 8. 68% were auto-generated templates, 22% had no links at all, and some files carried "blocking rules" the format has no power to enforce. If you've heard "add an llms.txt to get cited by AI search," this is worth reading before you build one.

Over the past two years you've probably run into the advice: "add an llms.txt file at your site's root so AI search can find you." The idea is a short file summarizing your site and listing its key pages, for AI systems to reference. Until now that advice has mostly been "people say it helps" — nobody had looked, at scale, at what these files actually contain in the wild.

On September 8, Common Crawl — the nonprofit that has spent over a decade collecting and publishing web-crawl data — did exactly that. It pulled every llms.txt file it had collected across the entire web: 584,107 of them.

What was actually inside

68% were auto-generated templates produced by website plugins, and 41% of those came from Wix specifically. 22% had no links at all — meaning the one job this file is supposed to do, list your key pages, was simply left blank. 6% carried "crawler instructions" like rate limits or copyright notices, and 32 of those files went further and explicitly tried to "block" Common Crawl's own bot. But those same sites' robots.txt files — the actual standard that crawlers are built to respect — contained no such blocking rule. Whatever you write in an llms.txt file has no real power to allow or block a crawler. Ten files even showed signs of prompt injection, text aimed at steering an AI into a specific answer.

A smaller minority did use it well — files generated by the All in One SEO plugin averaged a median of 138 links, functioning more like an actual sitemap.

The mechanism: not a permit, not a lock

llms.txt isn't a standard — it's a proposal anyone is free to ignore. Writing "don't use this page for AI" inside it binds no crawler to comply. No major AI service has ever officially stated it reads llms.txt. If you want to actually restrict access, that's what robots.txt is for — crawlers are built to respect it. If you want to be found, writing your site's content clearly does more work than any special AI-facing file. llms.txt sits somewhere in between: a note with no enforcement behind it.

What's confirmed and what isn't — The figures above come from Common Crawl's own published analysis (as of September 8), confirmed through Search Engine Journal's reporting. Common Crawl's original report page (on Hugging Face) was unreachable during this session's verification pass. What this analysis does NOT establish is whether llms.txt affects how often AI search actually cites a site — it only measured what's written inside these files, not whether writing them works.

If your site gets 100 visitors a week

Don't spend an evening building this file. There's no guarantee any crawler reads it, and skipping it costs you nothing measurable. Spend that time instead writing one more piece of content with real numbers and cases — the kind of thing AI search actually has something to cite. If you do add one, treat it as a short, human-readable summary page, not a special AI control panel. And if you actually want to control crawler access, remember that's robots.txt's job, not this file's.

Sources

Issue posts are researched and drafted by an automated AI pipeline, then published through editorial gates: a real event, linked sources, and a stated fact-check date.

Share this postXThreadsLinkedInHacker News

More issues