Royking Niba

Robots.txt for SEO: What It Controls, What It Cannot, and How Google Reads Yours

· Royking Niba

Stock photograph of rack-mounted networking equipment with ethernet cables plugged into its ports.

The single most common robots.txt mistake I find on a recovery audit is a site that has tried to remove pages from Google by adding a Disallow rule. It does not work, and Google’s own documentation says so in one sentence: “It is not a mechanism for keeping a web page out of Google. To keep a web page out of Google, block indexing with noindex or password-protect the page.”

The reason it does not work is mechanical. A crawl block stops Google reading the page, and a page Google cannot read is a page whose noindex tag Google also cannot read. What you get instead is the outcome Google describes: “Google can’t index the content of pages which are disallowed for crawling, but it may still index the URL and show it in search results without a snippet.” The URL stays in the index, stripped of everything that would have made it look legitimate.

This piece covers what the file is for, the decision table that maps each goal to the control that actually delivers it, and the documented behaviour that decides how Google reads your file at all.

What robots.txt is for

Google’s definition is narrow and worth taking literally: “A robots.txt file tells search engine crawlers which URLs the crawler can access on your site.” It is a crawl-traffic instrument. Google names two legitimate uses for it: “Use a robots.txt file to manage crawl traffic, and also to prevent image, video, and audio files from appearing in Google Search results.”

Note the asymmetry in that sentence. For media files, robots.txt does keep the file out of results. For an HTML page, it does not. That distinction is the whole source of the confusion, and it is the reason the same rule that correctly hides a PDF generation endpoint will leave a thin category page sitting in the index with no snippet.

Goal to control: the table I hand clients

Almost every robots.txt argument I have had with a developer resolves once we write the goal down and then ask which control delivers it. This is that table.

What you wantDoes robots.txt do it?The control that does
Keep an HTML page out of Google’s results entirelyNo. The URL can still be indexed without a snippetnoindex in a meta robots tag or X-Robots-Tag header, on a crawlable page
Keep a page behind a wall from anyone, Google includedNoAuthentication. Google names password protection as the reliable method
Stop an image, video or audio file appearing in resultsYes. This is a documented userobots.txt, or noimageindex where appropriate
Stop Googlebot wasting requests on infinite faceted URLsYes, and this is the best use of the filerobots.txt, with the parameter patterns disallowed
Consolidate two near-duplicate URLs into one indexed pageNo, and a Disallow actively prevents itA 301 redirect, or rel=canonical on a crawlable page
Remove a page from results urgentlyNoThe Removals tool in Search Console for a temporary block, then noindex or a 404 or 410
Stop a page’s link equity flowing onwardNo, and a disallowed page’s links are not read at allReconsider the requirement. This is almost never the real problem
Keep a staging site out of the indexNo. This is the classic failure caseHTTP authentication on the whole environment

Two rows in that table are the ones that cost money. Disallowing a page you were trying to canonicalise is self-defeating: Google needs to fetch the page to see the canonical link element, so the block guarantees the consolidation you wanted will not happen. And a staging site disallowed rather than authenticated is how duplicate copies of an entire site end up in the index, because a single external link to a staging URL is enough to get it indexed snippet-free.

How Google actually reads your file

The part of the specification that never makes it into checklists is what happens when the file is not a clean 200. The behaviour is documented and it decides whether your rules apply at all, which makes it a real diagnostic tool after an unexplained crawl change.

What your server returnsWhat Google does with itWhy it matters to you
2xx with a valid fileRules are parsed and appliedThe intended case
3xx redirectGoogle “follows at least five redirect hops as defined by RFC 1945 and then stops and treats it as a 404 for the robots.txt file”A redirect chain on robots.txt ends in your rules being ignored
4xx, other than 429Google’s crawlers “treat all 4xx errors, except 429, as if a valid robots.txt file didn’t exist”A 403 on robots.txt means everything is crawlable, not nothing
5xx server error“For the first 12 hours, Google stops crawling the site but keeps trying to fetch the robots.txt file. If Google can’t fetch a new version, for the next 30 days Google will use the last good version”A sustained 5xx on this one file can stop crawling of the whole site
A file larger than the limit“Google enforces a robots.txt file size limit of 500 kibibytes (KiB). Content which is after the maximum file size is ignored”Rules at the bottom of a generated 900 KiB file are not being read
Any change you just madeGoogle “generally caches the contents of robots.txt file for up to 24 hours, but may cache it longer in situations where refreshing the cached version isn’t possible”A fix is not instant, and neither is a mistake

Read the 4xx and 5xx rows next to each other, because they invert. A permissions error that hides your robots.txt opens the whole site to crawling. A server error on the same file closes it. If a client tells me crawling collapsed overnight with no deployment, the robots.txt response code is the first thing I check, and a 5xx on it explains the shape of the drop better than anything in the content.

The scope rule that catches people out

Google states it plainly: “The rules listed in the robots.txt file apply only to the host, protocol, and port number where the robots.txt file is hosted.” Three dimensions, all of which have to match, and each one is a real failure I have found on live sites.

A worked example. A site serves a careful robots.txt at https://www.example.com/robots.txt disallowing its internal search endpoint. That file governs exactly nothing on any of the following: https://example.com/ without the www, because that is a different host; http://www.example.com/, because that is a different protocol; and https://shop.example.com/, because a subdomain is a separate host and needs its own file. The canonical hostname usually redirects, which hides the gap. The subdomain frequently does not, which is how a shop or a blog subdomain ends up crawled under rules nobody wrote.

Google is also explicit about the limits of the mechanism itself, and two are worth remembering when someone asks you to block a scraper: not every crawler honours the file, and different crawlers interpret the syntax differently. A rule is a request, and a badly behaved bot ignores requests. Blocking those is a server-level job, not a robots.txt one.

One more trap: blocking resources

Google permits blocking unimportant resource files, with a condition attached: only where their absence will not make the page harder to understand. In practice that condition is failed more often than it is met. A theme that disallows /wp-content/ or a build pipeline that disallows a /static/ directory blocks the CSS and JavaScript the page needs to render, and Google is then assessing a page it cannot see properly. On a JavaScript-rendered site this is not a subtle degradation, it is the difference between a page with content and a page with none.

The robots.txt audit I run

  1. Fetch the file at every host, protocol and port combination the site answers on, including subdomains. Note which ones have no file at all.
  2. Check the HTTP status of each, not just the contents. A 403 or a redirect chain is a finding in itself.
  3. Check the file size against the 500 KiB limit. Generated files on large ecommerce sites do reach it.
  4. List every Disallow rule and, for each, write down the goal it was added for. Any rule whose goal is “keep this out of Google” is misfiled and needs a noindex instead, which means the rule has to come out first so the page can be crawled.
  5. Confirm no CSS, JavaScript or image path needed for rendering is disallowed.
  6. Cross-check the disallowed patterns against the Page Indexing report. URLs reported as indexed though blocked by robots.txt are exactly the failure this whole piece describes.
  7. Read the server logs for requests to disallowed paths. Googlebot honouring the rules and some other agent ignoring them is normal, and worth knowing before someone blames the file.

Step four is where the real work is. Removing a Disallow so that Google can crawl a page and read its noindex feels backwards to most developers, and it is the correct sequence. The page has to be readable to be told to leave.

Related reading

Sources: Google Search Central documentation, “Introduction to robots.txt” and “How Google interprets the robots.txt specification”, both checked 27 September 2026. Quoted sentences are Google’s own wording.

Leave a Reply

Your email address will not be published. Required fields are marked *