Unsourced
Analysis

llms.txt won't get you cited.

For two years the go-to AI-search tip has been the same: add an llms.txt file to your site and the assistants will read it and cite you. It is cheap, it is easy, and it is on every GEO checklist. The problem is what happens when you read the server logs. The crawlers it is written for are not fetching it, and publishing it does not move your citations. It is a courtesy file, not a citation strategy, and the evidence on that is not close.

Published 8 September 2026 · Unsourced Research · All research

The pitch, and then the logs

llms.txt is a small text file you place at the root of your site: a heading, a short summary, and a tidy list of links meant to hand an AI assistant the clean version of your content. Jeremy Howard proposed it in September 2024, and through 2026 it hardened into standard GEO advice. Add the file, the story goes, and the assistants will read it and be more likely to cite you.

It is a good story. It is just not what the traffic shows. When researchers stopped theorising and read the server logs, the crawlers the file is addressed to were not the ones fetching it.

Zero

An independent analysis tracked requests for llms.txt-family files across roughly 900 domains from September 2025 to April 2026. It logged 1,227 requests in total. The share that came from a verified frontier-lab crawler, the GPTBot, ClaudeBot, PerplexityBot and Google-Extended user-agents that actually drive AI answers, was zero.

The requests that did arrive were the mundane kind: commercial data aggregators (about two thirds), human browsers opening the file directly (about a third), and a handful of security scanners. Among the requesters there was not a single verified AI-lab bot. The file you are writing for the models is being read by everything except the models.

unsourced.app Who actually fetches llms.txt ~900 domains, 1,227 requests logged, Sept 2025 to Apr 2026 Data aggregators 65% Human browsers 32% Security scanners 3% Verified AI labs 0% None from a verified AI-lab crawler (GPTBot, ClaudeBot, PerplexityBot, Google-Extended).
Requests for llms.txt-family files by requester type, across ~900 monitored domains.

No lift, either

Maybe the file still helps indirectly. That was testable too. A separate study across roughly 300,000 domains, published in late 2025, checked whether publishing llms.txt correlated with more AI citations. It found no measurable link. In the researchers' framing the file "behaved as noise, not signal," and in one setup removing it slightly improved model accuracy. Two different methods, direct server logs and a large correlation study, land in the same place: no evidence of a citation benefit.

And no lab will confirm it

The obvious tiebreaker would be a statement from the labs themselves. There isn't one. No major AI company has said on record that its crawlers fetch and parse llms.txt, and none has explicitly denied it. So the file sits in the worst possible spot for a tactic: the intended readers neither endorse it nor demonstrably use it, and the traffic says they do not fetch it. You are optimising into a void.

Why this keeps happening

llms.txt is the cleanest example of a pattern that runs through AI-search advice. A low-effort action gets sold as a citation lever because it is easy to recommend and almost impossible for the person doing it to disprove. You add the file, nothing visibly breaks, and you have no way to check whether it did anything, so it stays on the checklist. The only thing that settles it is evidence, and the moment someone reads the logs, the tactic evaporates.

What to do instead

Stop writing files for readers who never arrive, and start measuring the readers who do. Two questions are answerable with evidence rather than hope. First, which AI crawlers actually reach your pages, confirmed against the operator's published IP ranges and reverse DNS instead of the user-agent they type, since the name alone is trivially spoofed. Second, whether any of those visits turns into a citation in a real answer. That is the whole game, and it is measurable.

This is what Unsourced does. We show which verified AI crawlers reached your pages, what they fetched, and whether it became a citation across the assistants we check, all backed by the answer and the named source rather than a checklist item you cannot verify. Keep the llms.txt file if you like, it costs nothing as a courtesy. Just do not confuse a courtesy with a strategy.

The takeaway

The uncomfortable rule of AI search is that most of what gets sold as a citation lever has no evidence behind it. llms.txt is the tidiest case: the file exists, the advice is everywhere, and the labs it is addressed to have fetched it, in the best data we have, zero times. Optimise for what you can prove.

Common questions

Does llms.txt help me get cited by AI?

The available evidence says no. A study across roughly 300,000 domains found no measurable link between publishing llms.txt and AI-citation frequency, and an independent server-log analysis of about 900 domains recorded zero requests for llms.txt files from any verified frontier-lab crawler. There is no measured citation lift to point to.

Do OpenAI, Anthropic or Perplexity actually read llms.txt?

None has confirmed on record that it fetches or parses the file, and none has denied it either. In server logs their crawlers were not among the requesters at all, so you would be optimising for a behaviour nobody has demonstrated.

Then who is fetching llms.txt?

In one analysis of about 900 domains the requests came from commercial data aggregators (roughly two thirds), human browsers (about a third) and security scanners (a few percent). Not one verified AI-lab crawler was among them.

Should I delete my llms.txt file?

No need. It is cheap to keep as a plain navigation aid and it does no harm. Just do not file it under citation strategy or expect it to change your AI visibility, because the data does not support that.

What should I do instead?

Optimise for what you can verify. See which AI crawlers actually reach your pages, confirmed by IP range and reverse DNS rather than the user-agent they claim, and whether any of it turns into a citation. Spend the effort on evidence, not on a file the labs are not reading.

Related

A crawl is not a citation

Being fetched and being named in an answer are different events. llms.txt is not even the first one.

Read it →

An AI visibility score is a guess

The same pattern: a number sold as insight, with no evidence underneath it.

Read it →

AI crawler directory

Which AI crawlers actually reach your site, how to verify them, and what each one does with your content.

Read it →

See the evidence for your own site

Unsourced captures which AI assistants cite you, proves which crawlers really fetched your pages, and re-checks after you act — evidence, not a score.

© Unsourced, the evidence layer for AI search. Our data and findings are free to quote and cite — please attribute to Unsourced and link to unsourced.app.