The llms.txt question, answered with the evidence
We publish an llms.txt file and we will tell you plainly that the evidence says it does almost nothing today. Here is the data, and here is why we ship it anyway.
There is a file called llms.txt that a lot of people are adding to their sites, on the theory that AI systems will read it and understand the site better.
We publish one. We also want to be straightforward about what the evidence says, because writing an article that implies it is essential would contradict everything else on this site.
What the data shows
Ahrefs analysed server-log data across 137,000 domains and looked at who was actually requesting llms.txt files.
97% of them received zero requests in the month studied.
Of the requests that did arrive, AI retrieval bots, the ones answering user queries inside AI search products, accounted for 1.1%. The rest were SEO audit tools, unidentified bots, ordinary web crawlers and technology profiling services.
Put plainly: the file is mostly being read by tools that check whether the file exists.
As of that study, none of the major AI providers had publicly stated that their production systems read or act on llms.txt.
Why we ship it anyway
Three reasons, none of which is “it will improve our AI visibility.”
It costs almost nothing. It is a text file generated from content that already exists. No ongoing maintenance, no risk, no trade-off against anything else.
It might matter later. Standards get adopted or abandoned, and the cost of having it if adoption happens is lower than the cost of not having it. That is an option, cheaply purchased, and it should be described as an option rather than as a strategy.
It is genuinely useful for a different reason. A well-written llms.txt is a clean, structured statement of what a site contains and what each page is for. Writing one forces a certain clarity about your own information architecture, and the file is a good thing to hand a person who asks what is on your site.
What we will not do is tell you it is a ranking factor, because we have no evidence that it is and there is evidence that it is not being read.
Why this note exists at all
Because we would look ridiculous doing otherwise.
This site has a page arguing that a gap is a finding and never a hole to fill, and that absence of evidence should be reported as absence of evidence rather than converted into a confident claim. Shipping an llms.txt and quietly implying it does something would be the site failing its own standard, on its own domain, in public.
There is also a wider point about the current state of advice around AI search optimisation. A great deal of it is confident, recent, and unsupported. Some of it is derived from how retrieval actually works. Most of it is inference from a handful of observations, presented with more certainty than the observations can carry.
The useful filter is the same one we apply to competitive research: what would this claim look like if it were false, and has anybody checked?
For llms.txt, somebody checked. 137,000 domains, server logs, 97% zero requests. That is a much stronger form of evidence than the case for it, which is mostly enthusiasm.
What the evidence does support
The tactics with real support behind them are dull, and their dullness is why they get less coverage.
Content that directly answers the question a person asked. Retrieval systems are matching a question to text. Text that contains a clear, complete answer to a real question is easier to retrieve and easier to quote.
Being cited by third parties. Corroboration from sources other than your own domain does substantial work. This is the single largest lever in most categories and it is also the slowest.
Structure that makes extraction easy. Clear headings, self-contained answers, definitions that stand alone. Not because a parser rewards markup, but because a self-contained paragraph is quotable and a paragraph that depends on three earlier paragraphs is not.
Original material. Something nobody else has, that other people reference. The compounding asset.
Every one of those is also just good publishing. That is not a coincidence, and it is a reasonable prior for evaluating any new tactic: if it only helps machines and not readers, be suspicious of the evidence.
Our own file
Ours is at /llms.txt. It lists the pages, says what each one covers, and states the things we will not do: no pricing, no unsourced numbers, no claims about capabilities that are designed rather than shipped.
If an AI system reads it, good. If not, it remains an accurate index of the site, which is worth having regardless of whether anything automated ever requests it.
That is the whole claim. We would rather make a small honest one than a large unsupported one, and this is a site about measurement, so the standard has to apply to us first.
The Ahrefs study, with its date and its method, is listed with every other figure we cite on the sources section of the methodology page.
Related
AI and evidence
Why we let a human overrule the adversarial checker
We measured what would happen if our skeptical second pass ran automatically. It would have demoted the strongest competitor most and handed our client back first place. A checker tuned to argue, applied mechanically, recreates the bias it was built to remove.
AI and evidence
We audited our own gate and it caught 2 of 7 fabrication classes
Quote verification is necessary and nowhere near sufficient. Here are the five fabrication classes it cannot see, why every one of them passes a byte-for-byte check, and what we built as a result.
AI and evidence
Negative controls for AI research
A verification step nobody tests is a verification step you are trusting on faith. Plant a deliberately false claim on every run, and fail loudly if it ever gets accepted.