Royking Niba

llms.txt: What the File Actually Does, Who Reads It, and What It Weighs on a Real Site

· Royking Niba

Stock photograph of colour-highlighted HTML code on a computer monitor. A generic illustration, not a screenshot of any llms.txt file or site discussed in the article.

llms.txt is a proposed plain-text file, written in Markdown and served at the root of a website, that gives AI agents a short, curated map of the pages worth reading. It was proposed in 2024 by Jeremy Howard, and its own specification says it is meant mainly for use at inference time, when an assistant is answering a question, rather than for model training. Google Search does not use it: Google’s own guide says creating one “will neither harm nor help” your visibility in Search. So it is cheap and optional, it is aimed at agents rather than search engines, and it is not an SEO ranking lever. This page shows what the specification actually requires, what Google has published about it, and what the file weighs on a real 40-post site.

What the specification actually says

The proposal lives at llmstxt.org. I read it on 3 October 2026. Its stated job is narrow:

A proposal to standardise on using an /llms.txt file to provide information to help agents use a website.

llmstxt.org, the /llms.txt specification

The format is stricter than most people assume, and looser in one place. The file must contain, in order: an optional byte-order mark, then “An H1 with the name of the project or site. This is the only required section”, then a blockquote with a short summary, then any number of Markdown sections that are not headings, then any number of sections “delimited by H2 headers” that list links. A section headed Optional has a defined meaning: it is “used, by convention, for secondary information: links an agent can skip when a shorter context is needed.”

Three details matter in practice:

  • Only the H1 is mandatory. A file with a title and nothing else is technically valid, and useless.
  • Location is flexible. “The file can be placed at the site root, or at any path within it, covering the pages under that path.” A documentation subfolder can carry its own.
  • It assumes Markdown copies of your pages. The proposal asks sites to “provide a clean markdown version of those pages at the same URL as the original page”, either as page.html.md or page.md. Most WordPress sites do not do this, and the file still works as a link list without it.

On purpose, the authors are direct: “Our expectation was that llms.txt would mainly be useful for inference rather than training.” That sentence alone answers most of the questions I get about it. It is not an opt-out from training and it is not an access control. For that you need robots.txt rules aimed at the specific crawlers, which I cover in how to get cited by ChatGPT.

What Google has published about it

Google’s AI optimization guide, last updated 10 July 2026, is the clearest primary source. Read on 3 October 2026:

You don’t need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search

Google Search Central, AI optimization guide

It’s completely fine if you decide to create and maintain LLMS.txt files (or other similar files) for other services or systems that use these files. Doing so will neither harm nor help your site’s visibility or rankings in Google Search

Google Search Central, AI optimization guide

The page on AI features in Search says the same thing from the eligibility side: “There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.” To be shown as a supporting link, “a page must be indexed and eligible to be shown in Google Search with a snippet.” That is the whole entry ticket. I unpack the structured-data version of the same claim in schema markup for AI search.

There is one place inside Google where the file does appear. Chrome’s Lighthouse has an experimental Agentic Browsing category, which its documentation says requires Chrome 150 or later, and one of its audits “Checks for the presence of a machine-readable summary at the domain root.” That is a readiness check for browser agents, not a Search signal. Search Engine Journal reported the split in May 2026, including John Mueller comparing llms.txt to the keywords meta tag. I have not seen Mueller’s remark in a Google document, so treat it as reported commentary; the AI optimization guide quoted above is the authoritative position.

llms.txt next to the files it gets confused with

Most of the bad advice I see comes from treating llms.txt as a fourth member of the robots.txt family. It is not. This table separates what each control does according to its own documentation.

File or controlWho it is forWhat it can doEffect in Google Search
robots.txtAny crawler that follows the Robots Exclusion ProtocolAsks crawlers not to fetch URLs. Can target named AI crawlersControls crawling, which I cover in robots.txt for SEO
XML sitemapSearch engine crawlersLists URLs you want discoveredA discovery hint, not a guarantee of crawling or indexing
nosnippet, max-snippet, data-nosnippetGoogle Search, including AI featuresLimits how much of a page can be shownGoogle lists these as the controls for limiting what its AI features show
llms.txtAI agents and tools that choose to read itSuggests which pages matter and in what orderNone. “Neither harm nor help”, in Google’s words

The practical reading: if you want to stop something, llms.txt is the wrong tool, because nothing obliges any system to read it. If you want to point a willing agent at your best pages, it is a reasonable, low-cost courtesy.

What the file weighs on a real site

Everyone argues about llms.txt in the abstract, so I measured it on this site. On 3 October 2026 roykingniba.com had 40 published posts. I pulled their titles, URLs, excerpts and body text straight from the database and counted characters. Token estimates use OpenAI’s published rule of thumb that “1 token is approximately 4 characters” and “approximately three-quarters of a word” for English, so treat them as estimates, not tokenizer output.

What you would publishCharactersEstimated tokensRelative size
llms.txt, one bare link line per post (title and URL)5,344about 1,3001x
llms.txt, link line plus the post excerpt as the note14,342about 3,6002.7x
Every post’s full text, tags stripped375,865about 85,000 to 94,00070x
Every post’s stored HTML503,359not estimated94x

The full text of 40 posts runs to 63,819 words. The excerpt-annotated index carries every one of those 40 URLs in about 4 percent of the characters. That ratio, roughly 26 to 1 between the full text and a useful index, is the honest case for the file: it lets an agent decide which two or three pages to fetch instead of reading everything. It also shows where the effort really goes. Writing 40 link lines is ten minutes of work. Writing 40 one-sentence notes that tell an agent what each page answers is the part that makes the file worth having, and it is the same skill as writing a good meta description.

The stored HTML is 34 percent heavier than the text inside it, which is the other half of the argument for Markdown copies of pages. Whether any given agent fetches those copies is not something the specification can promise.

A worked example

This is the shape I would ship for a consultant site like this one. Note the H2 sections, the one-line notes after each link, and the Optional section holding the pages an agent can drop when it is short of context.

# Royking Niba
> SEO and GEO consultant. Diagnoses traffic drops, Google penalties and spam-update hits, and fixes the technical causes. Every article cites Google's own documentation.
## Penalty and recovery
- [Google penalty recovery](https://roykingniba.com/google-penalty-recovery/): the five-step diagnosis before touching a single link
- [Manual action vs algorithmic penalty](https://roykingniba.com/manual-action-vs-algorithmic-penalty/): how to tell which one hit you
- [The disavow file, explained](https://roykingniba.com/disavow-file-explained/): when disavowing helps and when it hurts
## Technical SEO
- [Crawl budget](https://roykingniba.com/crawl-budget/): who actually has a problem
- [Index bloat](https://roykingniba.com/index-bloat/): the cost in recrawl days
## AI search
- [How to get cited by ChatGPT](https://roykingniba.com/how-to-get-cited-by-chatgpt/): the four OpenAI crawlers and the robots.txt lines that matter
## Optional
- [PBN backlinks](https://roykingniba.com/pbn-backlinks/): why every network decays

Two choices in that file are deliberate. The blockquote says what the site does in one sentence, because that is the line an agent is most likely to quote. And the Optional section holds evergreen background rather than the money pages, because the specification defines Optional as the part an agent may skip.

Should you bother

My answer depends on what the site is, not on what the file is.

  • Developer documentation, APIs and software products: yes. Agents acting on behalf of developers are the use case the proposal was written for, and a curated map saves them guessing.
  • Content and service sites: optional. Do it in an hour if you like the discipline of writing one-line page summaries, and expect no change in Google Search.
  • A site in a penalty or recovery cycle: not now. Nothing in an llms.txt file touches the cause of a traffic drop. Spend the hour on the diagnosis in Google penalty recovery instead.
  • Anyone hoping to block AI training: wrong file. Use robots.txt rules for the named crawlers.

The five checks I run on an existing llms.txt

  1. Fetch /llms.txt and confirm a 200 status with a plain-text or Markdown content type, not an HTML error page dressed as a 200.
  2. Confirm the first line is an H1. It is the only required section, and I still find files that open with a paragraph.
  3. Check every link returns 200 and points at an indexable, canonical URL. A link to a redirect or a noindexed page is a mixed message to any reader.
  4. Check robots.txt does not disallow the file itself, or the pages it lists, for the agents you want to reach.
  5. Check your server logs for requests to /llms.txt before you invest more in it. If nothing ever asks for the file, that is your answer about its value on your site. The method is in server log file analysis.

Related reading

Leave a Reply

Your email address will not be published. Required fields are marked *