RESEARCH · 6 MIN READ
llms.txt: what it does, what it doesn't, and whether to bother
llms.txt is a proposed convention: a markdown file at your domain root that gives AI systems a curated map of your most important pages. What it doesn't do is anything confirmed - no major answer engine has publicly committed to consuming it, adoption sits at roughly 5-10% per OrganiKPI's testing, and no independently measured citation lift exists. Our verdict: it costs 30 minutes, carries no risk, and might pay off. Do it, expect nothing.
What it actually is
The proposal, from Jeremy Howard in 2024, is simple: put a file at /llms.txt containing a plain-markdown summary of your site - what it is, and a curated list of your key pages with one-line descriptions. The idea borrows the shape of robots.txt but inverts the intent: robots.txt tells crawlers what to skip; llms.txt tells language models what matters. A companion pattern, llms-full.txt, inlines full page content for systems that want everything in one fetch.
What it is not
- Not a standard. No standards body has adopted it, and there is no spec compliance to fail.
- Not confirmed infrastructure. OpenAI, Anthropic, Google, and Perplexity have not publicly committed to consuming llms.txt in their answer products.
- Not access control. It grants and restricts nothing; robots.txt and crawler user-agent rules still do that job.
- Not measured. Per OrganiKPI's testing, adoption sits around 5-10% of sites, and there is no independently measured citation lift from adding one. We know of no study showing it moves visibility.
The honest case for doing it anyway
The cost is one static file and 30 minutes. The downside risk is zero - no engine penalizes you for having it. The upside is a lottery ticket: if any engine starts consuming it, early adopters get whatever benefit exists on day one, and some smaller agents and developer tools already fetch it opportunistically. There's also a side benefit that has nothing to do with machines: writing a one-page curated summary of your site is a useful forcing function. If you can't say what your ten most important pages are and why, that's worth discovering.
The honest case against expecting anything
The major engines already have crawling and retrieval pipelines built around HTML and search indexes; a parallel markdown channel solves a problem they may not feel they have. Adoption at 5-10% means models can't rely on it existing, which weakens the incentive to build consumption for it. And because no measurable lift has been demonstrated, any agency selling llms.txt creation as a visibility service is selling the 30-minute task, not a result. If it appears as a line item on a proposal, ask what evidence backs it.
If you do it, do it like this
- Keep it short and curated - your 10-20 most important pages, not a sitemap dump.
- Write the one-line descriptions to match your canonical entity description everywhere else. Consistency is the real signal; this file is one more place to reinforce it.
- Regenerate it when your site structure changes, or it becomes one more stale description working against you.
- Don't buy it as a service, and don't report it to anyone as an AI visibility win.
Where this sits in our own priority stack: below entity consistency, below fixing the pages behind your money prompts, below earning third-party mentions - and above nothing, because almost everything else costs more than 30 minutes. Calibrate accordingly.
Frequently asked
Does llms.txt actually improve AI visibility?
There is no evidence that it does. No major answer engine has publicly committed to consuming the file, and we know of no independently measured citation lift from adding one. Anyone selling llms.txt creation as a visibility service is selling a 30-minute task, not a result.
Is llms.txt the same as robots.txt?
No. They borrow the same shape and invert the intent. robots.txt tells crawlers what to skip and is honoured as an access convention; llms.txt tells language models which of your pages matter and grants or restricts nothing.
Should I create an llms.txt file?
Yes, but expect nothing from it. The cost is one static file and about 30 minutes, the downside risk is zero, and the upside is a lottery ticket if an engine starts consuming it. Keep it to your 10-20 most important pages and regenerate it when your site changes.
See where you stand.
Run a free scan and get your own answer-engine scorecard.