Royking Niba

Faceted Navigation SEO: The 4.5 Million URL Problem, and Which Control Actually Stops It

· Royking Niba

Stock photo of a person browsing shelves in a library archive, used here to illustrate filtering a large catalogue. It is not a screenshot of a website.

Faceted navigation is the most expensive feature most ecommerce sites ship without measuring. Every filter, sort and combination that produces a new URL is a new thing for a crawler to fetch, and the crawler cannot tell in advance that it is worthless. Google says so in plain language. The scale of it is arithmetic, and the arithmetic is worse than people expect.

Google’s own description of the problem

From Google’s documentation on crawling and managing faceted navigation, two sentences carry the whole case:

“Because the URLs created for the faceted navigation seem to be novel and crawlers can’t determine whether the URLs are going to be useful without crawling first, the crawlers will typically access a very large number of faceted navigation URLs before the crawlers’ processes determine the URLs are in fact useless.”

“If crawling is spent on useless URLs, the crawlers have less time to spend on new, useful URLs.”

Note what is not being claimed there. This is not a quality penalty and it is not a duplicate content sanction. It is a resource argument: crawling is finite, and filter permutations consume it. Google also warns that crawled faceted URLs mean “an increased resource usage on your server and, potentially, slower discovery of new URLs on your site”, which is the part your hosting bill notices before your rankings do.

The combinatorial arithmetic

Take a modest catalogue. Five facets, with 6, 8, 4, 5 and 3 values. Each facet can be unset, so each contributes its value count plus one. If every combination produces a crawlable URL:

StepCalculationURLs
Filter combinations in one category7 x 9 x 5 x 6 x 47,560
Across 40 categories7,560 x 40302,400
Times 3 sort orders302,400 x 3907,200
Times an average of 5 paginated pages each907,200 x 54,536,000
Crawlable URLs generated by a catalogue of a few thousand products4,536,000
Worked example. Facet counts are illustrative; multiply your own and the shape will not change.

That is the number Googlebot is being invited to work through in order to find a few thousand products. It is also why “we added a canonical, it’s fine” is not an answer on its own: a canonical is evaluated after the fetch, so the crawl has already happened.

The five controls, and what each one actually does

ControlStops the fetchRemoves it from the indexPasses link signals onwardThe catch
robots.txt disallow, e.g. disallow: /*?*products=YesNo, a disallowed URL can still be indexed from links aloneNoThe strongest crawl control and the bluntest. Block a parameter you later want crawled and you cannot see inside the page to find out
URL fragments for filter stateYes, because “Google Search generally doesn’t support URL fragments in crawling and indexing”Nothing was ever a separate URLNot applicableNeeds a front-end rebuild, and you give up the ability to have any filtered view ranked
rel="nofollow" on filter linksReduces discovery through those linksNoNoGoogle may still find the URL from a sitemap, an external link, or another internal path
rel="canonical" to the unfiltered pageNo. The fetch happens firstConsolidates, over timeYesGoogle notes it may decrease crawl volume of non-canonical versions over time. It is a hint, not a rule, and it is slow
noindexNo, the page must be crawled to read the tagYesLinks on a noindexed page are still followed initiallySolves an index problem and does nothing for the crawl problem in the short term

The column that matters is the first one. Only two controls stop the fetch, and only one of those is a configuration change rather than a rebuild. Everything else in that table is an indexing control being asked to do a crawling job.

When you do want filtered URLs crawled

Some filtered views have real demand behind them. A colour and size combination that people search for is a landing page, not waste. Google’s guidance for that case is specific:

  • “Use the industry standard URL parameter separator &.” Custom separators are not reliably parsed.
  • “Return an HTTP 404 status code when a filter combination doesn’t return results.” An empty results page returning 200 is a soft 404 factory, and it multiplies with every empty permutation.
  • Keep filter ordering in the URL consistent, so the same selection does not produce several different URLs.

My working rule on client builds: pick the handful of filter combinations that have verified search demand, give those clean static URLs with their own copy and internal links, and block the rest at the crawl layer. A curated dozen beats four and a half million every time, and it is the only version of this that a merchandising team can actually maintain.

How to size your own problem in an afternoon

  1. Multiply your facet value counts plus one, then multiply by categories, sort orders and average pagination depth. That is your theoretical ceiling.
  2. Open the Search Console crawl stats report and look at what share of requests hit URLs containing your filter parameters.
  3. Compare total known URLs against your actual product count. A ratio in the hundreds is the tell.
  4. Check what a filter combination with zero results returns. If it is 200 with an empty grid, fix that first, because it is cheap and it is generating soft 404s at scale.
  5. Decide which filtered views earn a crawl, then block the rest at the crawl layer rather than the index layer.

Related reading

Sources: Google Search Central, “Crawling and managing faceted navigation”, checked 28 September 2026.

Leave a Reply

Your email address will not be published. Required fields are marked *