X-Robots-Tag: The Header That Controls the 85 Percent of a Site a Robots Meta Tag Cannot Reach
· Royking Niba
The X-Robots-Tag is an HTTP response header that carries the same indexing and serving rules as a robots meta tag, with one decisive advantage: it works on files that cannot contain HTML. Google states the equivalence plainly, that "any rule that can be used in a robots meta tag can also be specified as an X-Robots-Tag", and then names the reason the header exists at all: "To block indexing of non-HTML resources, such as PDF files, video files, or image files, use the X-Robots-Tag response header instead." On most real sites that is the large majority of the indexable inventory.
The syntax, as Google documents it
From Google’s robots meta tag and X-Robots-Tag specification, checked on 2 October 2026, the header appears in the response alongside the status line and the other headers:
HTTP/1.1 200 OK
Date: Tue, 25 May 2010 21:42:43 GMT
(…)
X-Robots-Tag: noindex
(…)
Three mechanics follow from the same page. Directives can be stacked: "Multiple X-Robots-Tag headers can be combined within the HTTP response, or you can specify a comma-separated list of rules." The header can be aimed at one crawler, because it "may optionally specify a user agent before the rules", and Google supports exactly two user agent tokens in these rules, googlebot for all text results and googlebot-news for news results, with other values ignored. And when rules disagree, there is a stated winner: "In the case of conflicting robots rules, the more restrictive rule applies." Google’s own illustration is that a page carrying both max-snippet:50 and nosnippet ends up with nosnippet applied.
Every directive, and where each one can be delivered
The definitions in the second column are Google’s own wording from that specification. The third column is the point of this table: for a PDF, an image, a video file, a CSV or a JSON endpoint, the header is the only delivery method available, because there is no HTML head to put a meta tag in.
| Directive | What Google says it does | Works on non-HTML files |
|---|---|---|
all | "There are no restrictions for indexing or serving. This rule is the default value and has no effect if explicitly listed." | Header only |
noindex | "Do not show this page, media, or resource in search results." | Header only |
nofollow | "Do not follow the links on this page." | Header only |
none | "Equivalent to noindex, nofollow." | Header only |
nosnippet | "Do not show a text snippet or video preview in the search results for this page." | Header only |
indexifembedded | "Google is allowed to index the content of a page if it’s embedded in another page through iframes or similar HTML tags, in spite of a noindex rule." | Header only |
max-snippet:[number] | "Use a maximum of [number] characters as a textual snippet for this search result." | Header only |
max-image-preview:[setting] | "Set the maximum size of an image preview for this page in search results." | Header only |
max-video-preview:[number] | "Use a maximum of [number] seconds as a video snippet for videos on this page in search results." | Header only |
notranslate | "Don’t offer translation of this page in search results." | Header only |
noimageindex | "Do not index images on this page." | Header only |
unavailable_after:[date/time] | "Do not show this page in search results after the specified date/time." | Header only |
The original number: how much of a site the meta tag cannot reach
"Use the header for non-HTML files" sounds like an edge case until you count the files. An indexable resource is anything Google can return in a result, which includes images and PDFs, not only pages. These four inventories are illustrative profiles rather than measurements of any one site, and the shape they produce is consistent enough to be worth knowing before you plan any indexing work.
| Site profile | HTML pages | PDFs | Images | Video files | Total indexable resources | Reachable only by the header | Share |
|---|---|---|---|---|---|---|---|
| Small business site | 120 | 0 | 600 | 2 | 722 | 602 | 83.4% |
| Publisher archive | 8,000 | 400 | 36,000 | 600 | 45,000 | 37,000 | 82.2% |
| University or government site | 3,000 | 12,000 | 9,000 | 150 | 24,150 | 21,150 | 87.6% |
| Ecommerce catalogue | 14,000 | 250 | 98,000 | 400 | 112,650 | 98,650 | 87.6% |
| Across the four profiles | Header-only share | 82.2% to 87.6% |
Across all four profiles the robots meta tag can address between 12.4 and 17.8 percent of the indexable inventory. The header can address all of it. That is the practical case for knowing this header exists: an indexing policy expressed only in meta tags is a policy that covers roughly one resource in six, and the gap is usually filled with uploaded PDFs and media that nobody decided to make indexable in the first place.
One honest caveat on the numbers. Images are not indexed one-for-one the way pages are, and many of those 98,000 catalogue images are thumbnails and variants that would never rank on their own. The ratio overstates how many resources are worth a decision. It does not overstate the structural point, which is that the meta tag is unavailable for all of them regardless of how many matter.
The mistake that cancels the header silently
This is the failure I see most often, and it looks like good hygiene while it happens. A team wants a directory of PDFs out of search, so it does both of the things that sound right: it disallows the directory in robots.txt, and it sets X-Robots-Tag: noindex on the files. The two instructions cancel each other, because Google is explicit about the order of events:
If a page is disallowed from crawling through the robots.txt file, then any information about indexing or serving rules will not be found and will therefore be ignored.
Google Search Central, robots meta tag and X-Robots-Tag specifications
A disallowed file is never fetched, so its headers are never read, so the noindex is never seen. If those URLs were already indexed, or are linked from anywhere, they can stay in results indefinitely. The sequence that works is the opposite of what feels tidy: allow crawling, serve the noindex header, wait for the files to drop out of the index, and only then consider a robots.txt rule to stop the crawling. The table below is the version of this I hand to clients.
| What you did | Does Google fetch the file | Does Google see the header | Likely outcome |
|---|---|---|---|
| Header noindex, crawling allowed | Yes | Yes | The file drops out of search. This is the configuration you wanted. |
| Header noindex, directory disallowed in robots.txt | No | No | The rule is ignored. Already-indexed files can stay indexed. |
| Disallowed in robots.txt, no header | No | Not applicable | Crawling stops, indexing is not controlled. |
| Header carries both max-snippet:50 and nosnippet | Yes | Yes | nosnippet applies, because the more restrictive rule wins. |
| Header says noindex, page meta tag says index | Yes | Yes | noindex applies, by the same conflict rule. |
Where I actually use it
- Uploaded PDFs that duplicate a page’s content. The page should rank and the PDF should not, and
X-Robots-Tag: noindexon the PDF is the only way to say so. - Staging and preview hostnames, applied at the server or CDN for every response. One rule covers HTML, assets, feeds and downloads at once, which a meta tag deployment never does.
- Dated or expiring content, with
unavailable_after, where a listing genuinely stops being valid on a known date and nobody will remember to act on it. - Image-heavy sections where
noimageindexor amax-image-previewsetting is a licensing requirement rather than an SEO preference. - Non-HTML endpoints that crawl well and serve nobody: JSON and CSV exports, print views that render as files, generated calendar feeds.
One thing the header does not do, and it is worth saying because the two get conflated. A noindex header still costs you crawling. Google’s large-site crawl budget guide puts it directly: "Don’t use noindex, as Google will still request, but then drop the page when it sees a noindex meta tag or header in the HTTP response, wasting crawling time." If the goal is to remove a resource from search, the header is the right tool. If the goal is to stop Google spending requests on it at all, the header is not, and the guide’s answer to that is a 404, because "a 404 status code is a strong signal not to crawl that URL again".
How to check yours
- Request the resource and read the response headers, not the body.
curl -sI https://example.com/file.pdfis enough for one file. - Check the same URL as Googlebot, because the header can be set per user agent. A rule aimed at
googlebotwill not appear in a default request. - Confirm the URL is not disallowed in robots.txt before trusting any header you find on it.
- Crawl the non-HTML inventory specifically. Most site crawls are configured for pages and quietly skip the PDFs and media where this header matters most.
- Look for contradictions between the header and the page’s meta tag before changing either, remembering that the more restrictive rule is the one that will apply.
Leave a Reply