Faceted Navigation SEO: The 4.5 Million URL Problem, and Which Control Actually Stops It
· Royking Niba
Faceted navigation is the most expensive feature most ecommerce sites ship without measuring. Every filter, sort and combination that produces a new URL is a new thing for a crawler to fetch, and the crawler cannot tell in advance that it is worthless. Google says so in plain language. The scale of it is arithmetic, and the arithmetic is worse than people expect.
Google’s own description of the problem
From Google’s documentation on crawling and managing faceted navigation, two sentences carry the whole case:
“Because the URLs created for the faceted navigation seem to be novel and crawlers can’t determine whether the URLs are going to be useful without crawling first, the crawlers will typically access a very large number of faceted navigation URLs before the crawlers’ processes determine the URLs are in fact useless.”
“If crawling is spent on useless URLs, the crawlers have less time to spend on new, useful URLs.”
Note what is not being claimed there. This is not a quality penalty and it is not a duplicate content sanction. It is a resource argument: crawling is finite, and filter permutations consume it. Google also warns that crawled faceted URLs mean “an increased resource usage on your server and, potentially, slower discovery of new URLs on your site”, which is the part your hosting bill notices before your rankings do.
The combinatorial arithmetic
Take a modest catalogue. Five facets, with 6, 8, 4, 5 and 3 values. Each facet can be unset, so each contributes its value count plus one. If every combination produces a crawlable URL:
| Step | Calculation | URLs |
|---|---|---|
| Filter combinations in one category | 7 x 9 x 5 x 6 x 4 | 7,560 |
| Across 40 categories | 7,560 x 40 | 302,400 |
| Times 3 sort orders | 302,400 x 3 | 907,200 |
| Times an average of 5 paginated pages each | 907,200 x 5 | 4,536,000 |
| Crawlable URLs generated by a catalogue of a few thousand products | 4,536,000 | |
That is the number Googlebot is being invited to work through in order to find a few thousand products. It is also why “we added a canonical, it’s fine” is not an answer on its own: a canonical is evaluated after the fetch, so the crawl has already happened.
The five controls, and what each one actually does
| Control | Stops the fetch | Removes it from the index | Passes link signals onward | The catch |
|---|---|---|---|---|
robots.txt disallow, e.g. disallow: /*?*products= | Yes | No, a disallowed URL can still be indexed from links alone | No | The strongest crawl control and the bluntest. Block a parameter you later want crawled and you cannot see inside the page to find out |
| URL fragments for filter state | Yes, because “Google Search generally doesn’t support URL fragments in crawling and indexing” | Nothing was ever a separate URL | Not applicable | Needs a front-end rebuild, and you give up the ability to have any filtered view ranked |
rel="nofollow" on filter links | Reduces discovery through those links | No | No | Google may still find the URL from a sitemap, an external link, or another internal path |
rel="canonical" to the unfiltered page | No. The fetch happens first | Consolidates, over time | Yes | Google notes it may decrease crawl volume of non-canonical versions over time. It is a hint, not a rule, and it is slow |
noindex | No, the page must be crawled to read the tag | Yes | Links on a noindexed page are still followed initially | Solves an index problem and does nothing for the crawl problem in the short term |
The column that matters is the first one. Only two controls stop the fetch, and only one of those is a configuration change rather than a rebuild. Everything else in that table is an indexing control being asked to do a crawling job.
When you do want filtered URLs crawled
Some filtered views have real demand behind them. A colour and size combination that people search for is a landing page, not waste. Google’s guidance for that case is specific:
- “Use the industry standard URL parameter separator
&.” Custom separators are not reliably parsed. - “Return an HTTP 404 status code when a filter combination doesn’t return results.” An empty results page returning 200 is a soft 404 factory, and it multiplies with every empty permutation.
- Keep filter ordering in the URL consistent, so the same selection does not produce several different URLs.
My working rule on client builds: pick the handful of filter combinations that have verified search demand, give those clean static URLs with their own copy and internal links, and block the rest at the crawl layer. A curated dozen beats four and a half million every time, and it is the only version of this that a merchandising team can actually maintain.
How to size your own problem in an afternoon
- Multiply your facet value counts plus one, then multiply by categories, sort orders and average pagination depth. That is your theoretical ceiling.
- Open the Search Console crawl stats report and look at what share of requests hit URLs containing your filter parameters.
- Compare total known URLs against your actual product count. A ratio in the hundreds is the tell.
- Check what a filter combination with zero results returns. If it is 200 with an empty grid, fix that first, because it is cheap and it is generating soft 404s at scale.
- Decide which filtered views earn a crawl, then block the rest at the crawl layer rather than the index layer.
Related reading
- Google penalty recovery: the full diagnostic and recovery process
- Index bloat: what it costs in recrawl days, and the control that fixes it
- Pagination SEO: what Google actually crawls, and the 167-page example
- Crawl budget: who actually has a problem, and the index bloat behind it
- Robots.txt for SEO: what it controls, what it does not, and how Google handles failures
Sources: Google Search Central, “Crawling and managing faceted navigation”, checked 28 September 2026.
Leave a Reply