Robots.txt for SEO: What It Controls, What It Cannot, and How Google Reads Yours
· Royking Niba
The single most common robots.txt mistake I find on a recovery audit is a site that has tried to remove pages from Google by adding a Disallow rule. It does not work, and Google’s own documentation says so in one sentence: “It is not a mechanism for keeping a web page out of Google. To keep a web page out of Google, block indexing with noindex or password-protect the page.”
The reason it does not work is mechanical. A crawl block stops Google reading the page, and a page Google cannot read is a page whose noindex tag Google also cannot read. What you get instead is the outcome Google describes: “Google can’t index the content of pages which are disallowed for crawling, but it may still index the URL and show it in search results without a snippet.” The URL stays in the index, stripped of everything that would have made it look legitimate.
This piece covers what the file is for, the decision table that maps each goal to the control that actually delivers it, and the documented behaviour that decides how Google reads your file at all.
What robots.txt is for
Google’s definition is narrow and worth taking literally: “A robots.txt file tells search engine crawlers which URLs the crawler can access on your site.” It is a crawl-traffic instrument. Google names two legitimate uses for it: “Use a robots.txt file to manage crawl traffic, and also to prevent image, video, and audio files from appearing in Google Search results.”
Note the asymmetry in that sentence. For media files, robots.txt does keep the file out of results. For an HTML page, it does not. That distinction is the whole source of the confusion, and it is the reason the same rule that correctly hides a PDF generation endpoint will leave a thin category page sitting in the index with no snippet.
Goal to control: the table I hand clients
Almost every robots.txt argument I have had with a developer resolves once we write the goal down and then ask which control delivers it. This is that table.
| What you want | Does robots.txt do it? | The control that does |
|---|---|---|
| Keep an HTML page out of Google’s results entirely | No. The URL can still be indexed without a snippet | noindex in a meta robots tag or X-Robots-Tag header, on a crawlable page |
| Keep a page behind a wall from anyone, Google included | No | Authentication. Google names password protection as the reliable method |
| Stop an image, video or audio file appearing in results | Yes. This is a documented use | robots.txt, or noimageindex where appropriate |
| Stop Googlebot wasting requests on infinite faceted URLs | Yes, and this is the best use of the file | robots.txt, with the parameter patterns disallowed |
| Consolidate two near-duplicate URLs into one indexed page | No, and a Disallow actively prevents it | A 301 redirect, or rel=canonical on a crawlable page |
| Remove a page from results urgently | No | The Removals tool in Search Console for a temporary block, then noindex or a 404 or 410 |
| Stop a page’s link equity flowing onward | No, and a disallowed page’s links are not read at all | Reconsider the requirement. This is almost never the real problem |
| Keep a staging site out of the index | No. This is the classic failure case | HTTP authentication on the whole environment |
Two rows in that table are the ones that cost money. Disallowing a page you were trying to canonicalise is self-defeating: Google needs to fetch the page to see the canonical link element, so the block guarantees the consolidation you wanted will not happen. And a staging site disallowed rather than authenticated is how duplicate copies of an entire site end up in the index, because a single external link to a staging URL is enough to get it indexed snippet-free.
How Google actually reads your file
The part of the specification that never makes it into checklists is what happens when the file is not a clean 200. The behaviour is documented and it decides whether your rules apply at all, which makes it a real diagnostic tool after an unexplained crawl change.
| What your server returns | What Google does with it | Why it matters to you |
|---|---|---|
| 2xx with a valid file | Rules are parsed and applied | The intended case |
| 3xx redirect | Google “follows at least five redirect hops as defined by RFC 1945 and then stops and treats it as a 404 for the robots.txt file” | A redirect chain on robots.txt ends in your rules being ignored |
| 4xx, other than 429 | Google’s crawlers “treat all 4xx errors, except 429, as if a valid robots.txt file didn’t exist” | A 403 on robots.txt means everything is crawlable, not nothing |
| 5xx server error | “For the first 12 hours, Google stops crawling the site but keeps trying to fetch the robots.txt file. If Google can’t fetch a new version, for the next 30 days Google will use the last good version” | A sustained 5xx on this one file can stop crawling of the whole site |
| A file larger than the limit | “Google enforces a robots.txt file size limit of 500 kibibytes (KiB). Content which is after the maximum file size is ignored” | Rules at the bottom of a generated 900 KiB file are not being read |
| Any change you just made | Google “generally caches the contents of robots.txt file for up to 24 hours, but may cache it longer in situations where refreshing the cached version isn’t possible” | A fix is not instant, and neither is a mistake |
Read the 4xx and 5xx rows next to each other, because they invert. A permissions error that hides your robots.txt opens the whole site to crawling. A server error on the same file closes it. If a client tells me crawling collapsed overnight with no deployment, the robots.txt response code is the first thing I check, and a 5xx on it explains the shape of the drop better than anything in the content.
The scope rule that catches people out
Google states it plainly: “The rules listed in the robots.txt file apply only to the host, protocol, and port number where the robots.txt file is hosted.” Three dimensions, all of which have to match, and each one is a real failure I have found on live sites.
A worked example. A site serves a careful robots.txt at https://www.example.com/robots.txt disallowing its internal search endpoint. That file governs exactly nothing on any of the following: https://example.com/ without the www, because that is a different host; http://www.example.com/, because that is a different protocol; and https://shop.example.com/, because a subdomain is a separate host and needs its own file. The canonical hostname usually redirects, which hides the gap. The subdomain frequently does not, which is how a shop or a blog subdomain ends up crawled under rules nobody wrote.
Google is also explicit about the limits of the mechanism itself, and two are worth remembering when someone asks you to block a scraper: not every crawler honours the file, and different crawlers interpret the syntax differently. A rule is a request, and a badly behaved bot ignores requests. Blocking those is a server-level job, not a robots.txt one.
One more trap: blocking resources
Google permits blocking unimportant resource files, with a condition attached: only where their absence will not make the page harder to understand. In practice that condition is failed more often than it is met. A theme that disallows /wp-content/ or a build pipeline that disallows a /static/ directory blocks the CSS and JavaScript the page needs to render, and Google is then assessing a page it cannot see properly. On a JavaScript-rendered site this is not a subtle degradation, it is the difference between a page with content and a page with none.
The robots.txt audit I run
- Fetch the file at every host, protocol and port combination the site answers on, including subdomains. Note which ones have no file at all.
- Check the HTTP status of each, not just the contents. A 403 or a redirect chain is a finding in itself.
- Check the file size against the 500 KiB limit. Generated files on large ecommerce sites do reach it.
- List every
Disallowrule and, for each, write down the goal it was added for. Any rule whose goal is “keep this out of Google” is misfiled and needs anoindexinstead, which means the rule has to come out first so the page can be crawled. - Confirm no CSS, JavaScript or image path needed for rendering is disallowed.
- Cross-check the disallowed patterns against the Page Indexing report. URLs reported as indexed though blocked by robots.txt are exactly the failure this whole piece describes.
- Read the server logs for requests to disallowed paths. Googlebot honouring the rules and some other agent ignoring them is normal, and worth knowing before someone blames the file.
Step four is where the real work is. Removing a Disallow so that Google can crawl a page and read its noindex feels backwards to most developers, and it is the correct sequence. The page has to be readable to be told to leave.
Related reading
- X-Robots-Tag: the header that controls what a robots meta tag cannot reach
- JavaScript SEO: the three phases Googlebot uses, because a Disallow on your script bundle stops the page rendering at all.
- Orphan pages: why Google never finds them, because a Disallow rule and a noindexed archive both remove discovery routes.
- Faceted navigation SEO: the 4.5 million URL problem, and which control actually stops it
- Crawl budget, where disallowing junk URL patterns is the one job robots.txt does better than anything else.
- How to get cited by ChatGPT, which is the same file governing a different set of crawlers.
- Canonical tag SEO, and why a Disallow on a duplicate destroys the consolidation you were trying to achieve.
- Soft 404s, the other place where the HTTP response and the human-visible page say different things.
- Google penalty recovery, where a crawl-level misconfiguration gets misdiagnosed as a penalty.
Sources: Google Search Central documentation, “Introduction to robots.txt” and “How Google interprets the robots.txt specification”, both checked 27 September 2026. Quoted sentences are Google’s own wording.
Leave a Reply