Crawl Errors: What Google Does With Each Response Code, and the One It Calls “Possibly Good”
· Royking Niba
Most crawl error reports get read as a list of things to fix, top to bottom. Google does not treat them that way, and its own documentation says so: in the Crawl Stats report a 404 is filed under “Possibly good”, while a persistent 5xx is the one response that gets URLs taken out of the index. Below is what the documentation actually commits to for each status code, a triage order built from that, and the redirect arithmetic that explains why a site with no errors at all can still be wasting three quarters of its crawl.
What Googlebot does with each response, in Google’s words
Quoted from Google Search Central’s documentation on HTTP status codes, network and DNS errors, last updated 4 February 2026, checked today.
Google considers the content for processing (for example, in the case of Google Search, for indexing).
On 2xx responses
By default, Google’s crawlers follow up to 10 redirect hops.
On 3xx responses
Google doesn’t use the content from URLs that return 4xx status codes.
On 4xx responses
Google’s crawlers treat the 429 status code as a signal that the server is overloaded, and it’s considered a server error.
On 429
For Google Search, Google’s indexing pipeline removes from the index URLs that persistently return a server error.
On 5xx responses
Four of those five sentences describe something recoverable. One does not. “Removes from the index” is the only outcome on the list that destroys something you then have to earn back, and it is attached to the server-error family, which includes 429. That single asymmetry is the whole triage order.
The triage table
The middle column is quoted or paraphrased from the documentation above and from the Crawl Stats report help page. The severity ranking in the right-hand column is mine, and it is ordered by what the error destroys rather than by how many rows of it you have.
| Response | What Google documents | What it costs | Fix order |
|---|---|---|---|
| 5xx, persistent | Indexing pipeline removes the URL from the index | Lost indexing you have to re-earn | 1, same day |
| 429 | Treated as a server error | Same exposure as 5xx, often self-inflicted by rate limiting | 1, same day |
| DNS and connectivity failures | A Crawl Stats host availability category alongside robots.txt fetching | Sitewide, because nothing can be fetched | 1, same day |
| robots.txt unreachable | A Crawl Stats host availability category in its own right | Sitewide crawl behaviour, not one URL | 2 |
| Redirect errors and chains over the limit | Crawlers follow up to 10 hops by default | Wasted crawl, and a dead end past the limit | 3 |
| 401 and 407 | Filed under “Bad” in Crawl Stats | Content is unreachable, usually a staging leak or a misapplied gate | 3 |
| 4xx other than 401 and 407 | Google does not use the content | Nothing, if the URL is genuinely gone | 4 |
| 404 | Filed under “Possibly good” in Crawl Stats | Usually nothing at all | Last, and often never |
| Eight rows | Three are same-day | One row, 404, is the one most reports put first |
Google’s own classification is the point here. From the Crawl Stats report documentation, responses are grouped as good, which covers 200, 301, 302 and 304, possibly good, which is 404 on its own, and bad, which covers 401 and 407, 5xx, DNS issues, fetch errors, timeouts and redirect errors. A 404 for a page that no longer exists is the correct response and Google says as much by not calling it bad. Chasing those to zero is work that buys nothing.
Host status, which is the report most people never open
The Crawl Stats report assesses host availability in three categories: robots.txt fetching, DNS resolution and server connectivity. It shows three states, no issues, issues more than a week ago, and issues within the last week, over a rolling ninety-day window.
That amber middle state is the useful one and the one nobody checks, because it is the only place a crawl outage from three weeks ago leaves a visible mark after the graphs have recovered. If a client’s traffic stepped down and nobody can say why, host status is a thirty-second check that either rules a crawl outage in or out.
Google also scopes who this is for, which is worth quoting because it saves a lot of wasted effort on small sites:
This report is aimed at advanced users. If you have a site with fewer than a thousand pages, you should not need to use this report or worry about this level of crawling detail.
The report is also only available on root-level properties, which is why it appears empty for anyone who set Search Console up on a subfolder.
The worked example: where three quarters of a crawl goes
Redirects are not errors and they do not show up as errors, which is exactly why they are the biggest crawl leak on most migrated sites. The documented limit is ten hops, so a chain of three is well inside the rules and still costs you every time.
Take 2,000 URLs that are linked internally through a three-hop chain, which is the normal result of two replatforms and a protocol change nobody ever flattened.
| Chain length | Requests per URL | Total requests for 2,000 URLs | Requests returning no content | Share of crawl wasted |
|---|---|---|---|---|
| 3 hops | 4 (three redirects plus the document) | 8,000 | 6,000 | 75.0% |
| 2 hops | 3 | 6,000 | 4,000 | 66.7% |
| 1 hop | 2 | 4,000 | 2,000 | 50.0% |
| 0 hops, links updated | 1 | 2,000 | 0 | 0.0% |
| 3 hops to 0 hops | 8,000 down to 2,000 | 6,000 requests saved | 75% reduction |
Updating the internal links rather than leaving the redirects to absorb the traffic cuts the request count for that set of pages by three quarters. Nothing in the error report changes, because none of this was ever an error. A chain of eleven hops, by contrast, is a dead end under the documented ten-hop default, and that one does break.
What this model assumes
It assumes every one of the 2,000 URLs is actually fetched, that each chain is the same length, and that Googlebot does not cache an intermediate hop across requests. Real chains vary in length and real crawlers do cache, so treat 75 percent as the ceiling on the saving rather than a forecast. The conclusion survives the caveats: redirect chains are invisible to the error report and are usually the largest recoverable waste in the crawl.
The order I work in
- Check host status first, including the amber state, before reading a single URL-level row.
- Clear anything in the server-error family on the same day, 429 included, because that family is the only one Google documents as removing URLs from the index.
- Confirm robots.txt is fetchable, since its availability changes crawl behaviour sitewide rather than for one URL.
- Flatten redirect chains and update the internal links that feed them, then re-measure the request count rather than the error count.
- Leave 404s alone unless the URL should exist, has links pointing at it, or used to rank. Google files them as possibly good and it is right to.
Search Console tells you what Google concluded about a response. Only the server log tells you what Googlebot actually requested and in what order, which is why the redirect waste above shows up in logs long before it shows up anywhere else.
Related reading
- Google Penalty Recovery: how I diagnose and reverse a traffic collapse
- Server log file analysis for SEO: the only record of what Googlebot actually did
- Soft 404s: what Google actually means, and the correct fix for each cause
- 301 vs 302 redirect: which one Google canonicalizes, and what the wrong choice costs
- Robots.txt for SEO: what it controls, what it cannot, and how Google reads yours
- Page experience: Google says there is no single signal, and the 75th percentile is why that matters
Leave a Reply