Index Bloat: What It Costs in Recrawl Days, and the Control That Actually Fixes It
· Royking Niba
Index bloat is the gap between the number of URLs on a site that can be indexed and the number that are worth indexing. It is not a Google term and there is no report that names it, which is exactly why it goes unmeasured. It is measurable, though, and the cost is not abstract: every extra indexable URL competes for the same finite crawling, and the arithmetic below shows how quickly that pushes the pages you care about to the back of the queue.
What Google actually says about the cost
Google does not use the phrase, but it documents the mechanism plainly on its large-site crawling page, checked on 30 September 2026:
If Google spends too much time crawling URLs that it shouldn’t, Google’s crawlers might not explore the rest of your site.
Google Search Central, Large site owner’s guide to managing your crawl budget
The same page instructs you to "eliminate duplicate content to focus crawling on unique content rather than unique URLs", warns that "soft 404 pages will continue to be crawled, and waste your budget", and is explicit that noindex is the wrong instrument for the job: "Don’t use noindex, as Google will still request, but then drop the page, wasting crawling time." It also notes that "a 404 status code is a strong signal not to crawl that URL again."
That last pair is the part most cleanup projects get backwards. Adding noindex to twenty thousand thin URLs removes them from the index and leaves the crawling cost exactly where it was.
The original number: what bloat costs in recrawl days
Take a site with 4,000 URLs that genuinely earn their place, and assume Googlebot requests roughly 5,000 URLs a day from it. If crawling is distributed across the indexable set rather than concentrated on the valuable part, the share of daily requests landing on a page that matters is simply the ratio of valuable URLs to total indexable URLs. The final column is how long it then takes for every valuable page to be requested once.
| Total indexable URLs | Bloat ratio | Share of crawling on useful URLs | Useful fetches per day | Days to recrawl all 4,000 |
|---|---|---|---|---|
| 4,000 | 1.0x | 100.0% | 5,000 | 0.8 |
| 8,000 | 2.0x | 50.0% | 2,500 | 1.6 |
| 12,500 | 3.1x | 32.0% | 1,600 | 2.5 |
| 20,000 | 5.0x | 20.0% | 1,000 | 4.0 |
| 50,000 | 12.5x | 8.0% | 400 | 10.0 |
| 100,000 | 25.0x | 4.0% | 200 | 20.0 |
The shape is the point. Going from no bloat to a 5x ratio does not cost you a fifth of your crawling, it multiplies your recrawl cycle from under a day to four days. At 12.5x, a price change or a new article waits ten days on average to be seen. A publisher whose whole proposition is freshness has, at that ratio, given away freshness without changing a single article.
The model is deliberately simple and it overstates the effect at the low end, because Google does not crawl uniformly: it concentrates on URLs it considers valuable and updates frequently, and demand "varies based on a site’s size, update frequency, page quality, and relevance, compared to other sites". Treat the table as the direction and the magnitude, not a forecast. The useful exercise is to substitute your own two numbers from Search Console and your own indexable count, and see which row you are on.
Measuring your own ratio in four numbers
- Count indexable URLs: crawl the site and count URLs returning 200 with no noindex and a self-referencing canonical. This is the denominator and it is almost always larger than anyone expects.
- Count valuable URLs: the pages that could plausibly answer a search, receive a link, or make money. Templates, tag archives, empty filters, session variants and paginated tails of dead categories are not on this list.
- Take the daily crawl request figure from the Search Console crawl stats report.
- Divide. The ratio and the recrawl figure are the two numbers to take to whoever owns the templates.
Where the extra URLs come from, in order of how often they do
- Filter and sort parameters that produce a distinct URL without changing what the page says.
- Internal search result pages that are linked or sitemapped.
- Tag and archive taxonomies created one term at a time, most holding a single post.
- Paginated sequences with no end condition, still serving pages beyond the last item.
- Session, tracking and campaign parameters appended to otherwise clean URLs.
- Staging, print and AMP-era duplicates that were never removed.
- Soft 404s, which Google says "will continue to be crawled" while returning nothing of value.
Choosing the control, and why noindex is usually the wrong one
Each control does one job, and the documentation is precise about which. The question to ask is not "how do I get this out of the index" but "do I want this URL to stop existing, stop being fetched, or stop being a separate thing".
- The URL should not exist. Return 410, or 404. Google states that a 404 "is a strong signal not to crawl that URL again", so this is the only control that reduces both indexing and crawling permanently.
- The URL should exist for users but not be fetched. Disallow it in robots.txt. Note the trade: a disallowed URL can still be listed if other pages link to it, because Googlebot never reads the page to find any rule inside it.
- The URL is a variant of another page. Redirect it, or canonicalise it. This consolidates rather than removes, and it keeps whatever signals the variant had collected.
- The URL must stay live, be crawlable, and stay out of the index. Only then is noindex the right instrument, and it must remain crawlable to work: "For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler."
The combination to avoid is noindex plus a robots.txt disallow on the same URL. Googlebot cannot read a rule on a page it is not allowed to fetch, so the page keeps its chance of appearing in results while you believe it has been removed. It is the single most common self-inflicted wound in an index cleanup.
Sequencing a cleanup without breaking anything
- Start with the URLs that should not exist at all and return 410 for them. This is the only step that buys back crawling.
- Consolidate variants next, with redirects where a single successor exists and canonicals where it does not.
- Apply robots.txt disallow to parameter patterns that generate URLs faster than you can remove them, after confirming none of those URLs currently receive traffic or links.
- Reserve noindex for the small set that must stay live and crawlable.
- Stop the source: fix the template or the parameter handling that generated the URLs, or the count returns within a quarter.
- Remeasure the ratio after four weeks and compare the recrawl figure, which is the number that shows whether the work paid.
Related reading
- Faceted navigation, and the combinatorial arithmetic behind it
- Canonical tags: what overrules them and why
- Discovered, currently not indexed: a capacity backlog or a demand problem
- Duplicate without user-selected canonical: sampling the report
- Crawled, currently not indexed: finding the template behind it
- Alternate page with proper canonical tag: the three checks that decide whether to act
- Soft 404s, and what Google does with them
- Google penalty recovery: the full process
Leave a Reply