Royking Niba

X-Robots-Tag: The Header That Controls the 85 Percent of a Site a Robots Meta Tag Cannot Reach

· Royking Niba

Stock photograph of ethernet cables plugged into a network switch panel. It is a generic stock image and not a photograph of any particular server or client infrastructure.

The X-Robots-Tag is an HTTP response header that carries the same indexing and serving rules as a robots meta tag, with one decisive advantage: it works on files that cannot contain HTML. Google states the equivalence plainly, that "any rule that can be used in a robots meta tag can also be specified as an X-Robots-Tag", and then names the reason the header exists at all: "To block indexing of non-HTML resources, such as PDF files, video files, or image files, use the X-Robots-Tag response header instead." On most real sites that is the large majority of the indexable inventory.

The syntax, as Google documents it

From Google’s robots meta tag and X-Robots-Tag specification, checked on 2 October 2026, the header appears in the response alongside the status line and the other headers:

HTTP/1.1 200 OK
Date: Tue, 25 May 2010 21:42:43 GMT
(…)
X-Robots-Tag: noindex
(…)

Three mechanics follow from the same page. Directives can be stacked: "Multiple X-Robots-Tag headers can be combined within the HTTP response, or you can specify a comma-separated list of rules." The header can be aimed at one crawler, because it "may optionally specify a user agent before the rules", and Google supports exactly two user agent tokens in these rules, googlebot for all text results and googlebot-news for news results, with other values ignored. And when rules disagree, there is a stated winner: "In the case of conflicting robots rules, the more restrictive rule applies." Google’s own illustration is that a page carrying both max-snippet:50 and nosnippet ends up with nosnippet applied.

Every directive, and where each one can be delivered

The definitions in the second column are Google’s own wording from that specification. The third column is the point of this table: for a PDF, an image, a video file, a CSV or a JSON endpoint, the header is the only delivery method available, because there is no HTML head to put a meta tag in.

DirectiveWhat Google says it doesWorks on non-HTML files
all"There are no restrictions for indexing or serving. This rule is the default value and has no effect if explicitly listed."Header only
noindex"Do not show this page, media, or resource in search results."Header only
nofollow"Do not follow the links on this page."Header only
none"Equivalent to noindex, nofollow."Header only
nosnippet"Do not show a text snippet or video preview in the search results for this page."Header only
indexifembedded"Google is allowed to index the content of a page if it’s embedded in another page through iframes or similar HTML tags, in spite of a noindex rule."Header only
max-snippet:[number]"Use a maximum of [number] characters as a textual snippet for this search result."Header only
max-image-preview:[setting]"Set the maximum size of an image preview for this page in search results."Header only
max-video-preview:[number]"Use a maximum of [number] seconds as a video snippet for videos on this page in search results."Header only
notranslate"Don’t offer translation of this page in search results."Header only
noimageindex"Do not index images on this page."Header only
unavailable_after:[date/time]"Do not show this page in search results after the specified date/time."Header only
Definitions quoted from Google’s robots meta tag, data-nosnippet, and X-Robots-Tag specifications page, checked 2 October 2026. “Header only” means the meta tag is not an option for that file type, not that the meta tag cannot carry the rule on an HTML page.

The original number: how much of a site the meta tag cannot reach

"Use the header for non-HTML files" sounds like an edge case until you count the files. An indexable resource is anything Google can return in a result, which includes images and PDFs, not only pages. These four inventories are illustrative profiles rather than measurements of any one site, and the shape they produce is consistent enough to be worth knowing before you plan any indexing work.

Site profileHTML pagesPDFsImagesVideo filesTotal indexable resourcesReachable only by the headerShare
Small business site1200600272260283.4%
Publisher archive8,00040036,00060045,00037,00082.2%
University or government site3,00012,0009,00015024,15021,15087.6%
Ecommerce catalogue14,00025098,000400112,65098,65087.6%
Across the four profilesHeader-only share82.2% to 87.6%
Illustrative inventories, not audited sites. The counts are the kind of ratios these site types produce, and the figures are given so the arithmetic can be checked and replaced with your own.

Across all four profiles the robots meta tag can address between 12.4 and 17.8 percent of the indexable inventory. The header can address all of it. That is the practical case for knowing this header exists: an indexing policy expressed only in meta tags is a policy that covers roughly one resource in six, and the gap is usually filled with uploaded PDFs and media that nobody decided to make indexable in the first place.

One honest caveat on the numbers. Images are not indexed one-for-one the way pages are, and many of those 98,000 catalogue images are thumbnails and variants that would never rank on their own. The ratio overstates how many resources are worth a decision. It does not overstate the structural point, which is that the meta tag is unavailable for all of them regardless of how many matter.

The mistake that cancels the header silently

This is the failure I see most often, and it looks like good hygiene while it happens. A team wants a directory of PDFs out of search, so it does both of the things that sound right: it disallows the directory in robots.txt, and it sets X-Robots-Tag: noindex on the files. The two instructions cancel each other, because Google is explicit about the order of events:

If a page is disallowed from crawling through the robots.txt file, then any information about indexing or serving rules will not be found and will therefore be ignored.

Google Search Central, robots meta tag and X-Robots-Tag specifications

A disallowed file is never fetched, so its headers are never read, so the noindex is never seen. If those URLs were already indexed, or are linked from anywhere, they can stay in results indefinitely. The sequence that works is the opposite of what feels tidy: allow crawling, serve the noindex header, wait for the files to drop out of the index, and only then consider a robots.txt rule to stop the crawling. The table below is the version of this I hand to clients.

What you didDoes Google fetch the fileDoes Google see the headerLikely outcome
Header noindex, crawling allowedYesYesThe file drops out of search. This is the configuration you wanted.
Header noindex, directory disallowed in robots.txtNoNoThe rule is ignored. Already-indexed files can stay indexed.
Disallowed in robots.txt, no headerNoNot applicableCrawling stops, indexing is not controlled.
Header carries both max-snippet:50 and nosnippetYesYesnosnippet applies, because the more restrictive rule wins.
Header says noindex, page meta tag says indexYesYesnoindex applies, by the same conflict rule.
Rows one to three follow from Google’s statement that rules on an uncrawled URL are not found and are ignored. Rows four and five follow from its conflicting-rules statement and its own max-snippet example.

Where I actually use it

  • Uploaded PDFs that duplicate a page’s content. The page should rank and the PDF should not, and X-Robots-Tag: noindex on the PDF is the only way to say so.
  • Staging and preview hostnames, applied at the server or CDN for every response. One rule covers HTML, assets, feeds and downloads at once, which a meta tag deployment never does.
  • Dated or expiring content, with unavailable_after, where a listing genuinely stops being valid on a known date and nobody will remember to act on it.
  • Image-heavy sections where noimageindex or a max-image-preview setting is a licensing requirement rather than an SEO preference.
  • Non-HTML endpoints that crawl well and serve nobody: JSON and CSV exports, print views that render as files, generated calendar feeds.

One thing the header does not do, and it is worth saying because the two get conflated. A noindex header still costs you crawling. Google’s large-site crawl budget guide puts it directly: "Don’t use noindex, as Google will still request, but then drop the page when it sees a noindex meta tag or header in the HTTP response, wasting crawling time." If the goal is to remove a resource from search, the header is the right tool. If the goal is to stop Google spending requests on it at all, the header is not, and the guide’s answer to that is a 404, because "a 404 status code is a strong signal not to crawl that URL again".

How to check yours

  1. Request the resource and read the response headers, not the body. curl -sI https://example.com/file.pdf is enough for one file.
  2. Check the same URL as Googlebot, because the header can be set per user agent. A rule aimed at googlebot will not appear in a default request.
  3. Confirm the URL is not disallowed in robots.txt before trusting any header you find on it.
  4. Crawl the non-HTML inventory specifically. Most site crawls are configured for pages and quietly skip the PDFs and media where this header matters most.
  5. Look for contradictions between the header and the page’s meta tag before changing either, remembering that the more restrictive rule is the one that will apply.

Related reading

Leave a Reply

Your email address will not be published. Required fields are marked *