Royking Niba

Site Architecture for SEO: The Depth Arithmetic That Decides What Gets Crawled

· Royking Niba

Stock photograph of networking equipment with ethernet cables plugged into it. It illustrates the idea of a link graph and is not a photograph of a client system.

Site architecture is the link graph, not the folder names. Google reaches a page by following an anchor element from a page it already knows, so the question that decides whether your pages get crawled is not how tidy your URLs look but how many links stand between the homepage and the page you care about, and whether those links are ones Google can follow at all. This piece sets out what Google actually documents about crawlable links and URL structure, then gives you the arithmetic that tells you how deep your site is allowed to be.

What Google documents, and what it does not

Google publishes very little about "site architecture" under that name. It publishes a great deal about the two mechanisms architecture is made of: links that can be followed, and URLs that do not multiply. Those two pages are the ones to design against.

Google can only crawl your link if it’s an <a> HTML element (also known as anchor element) with an href attribute.

Google Search Central, Link best practices for Google, checked 30 September 2026

The same page adds that links inserted by JavaScript are fine, on one condition: "Links are also crawlable when you use JavaScript to insert them into a page dynamically as long as it uses the HTML markup shown above." The failure mode is not JavaScript. The failure mode is a clickable thing that is not an anchor with an href, such as a span with an onclick handler or a framework directive that never renders an href. Those are navigation to a user and invisible to a crawler, and a whole section of a site can sit behind them.

On URLs, the guidance is about multiplication rather than aesthetics:

Overly complex URLs, especially those containing multiple parameters, can cause problems for crawlers by creating unnecessarily high numbers of URLs that point to identical or similar content on your site.

Google Search Central, Keep a simple URL structure, checked 30 September 2026

Two concrete rules sit alongside it. Use readable words rather than long ID numbers. Use hyphens rather than underscores to separate words, "as it helps users and search engines better identify concepts". And trim parameters that do not change the content. None of that is a ranking trick. All of it reduces the number of distinct URLs a crawler has to work through to see your content once.

The original number: how deep a site is allowed to be

Depth is usually argued about in slogans, the three-click rule being the most common. It is arithmetic, and the arithmetic is unforgiving in a useful way. If every page carries the same number of outgoing internal links to pages not yet reached, the number of pages reachable at each click depth is that branching factor raised to the depth. The table below is computed from that model, and the cumulative column is the total number of URLs reachable at that depth or shallower.

Internal links per pageDepth 1Depth 2Depth 3Depth 4Cumulative at depth 4
10101001,00010,00011,111
20204008,000160,000168,421
303090027,000810,000837,931
50502,500125,0006,250,0006,377,551
10010010,0001,000,000100,000,000101,010,101
Pages reachable by click depth, computed from a uniform branching model. Real sites overlap heavily, so treat these as ceilings, not forecasts.

Read the row that matches your navigation. A site with 30 genuine internal links on an average page can in principle put 27,000 URLs within three clicks of the homepage, and 810,000 within four. A 50,000-URL catalogue therefore has no arithmetic excuse for being eight clicks deep. If it is, the cause is not size. The cause is that most of those 30 links point at the same handful of pages: the header menu, the footer, the same five category links on every template. Overlap is what collapses a theoretical ceiling into a real site where half the inventory is nine clicks down.

That is the practical value of the model. It gives you a number to test against. Crawl your own site, take the median count of distinct internal link targets per page, look up the row, and compare the ceiling with the depth your crawler actually reports. The gap between them is the amount of your navigation that is duplicated rather than distributive.

The second number: a crawlable-link compliance matrix

Depth only counts if the links are ones Google follows. This matrix maps the patterns that appear in real templates against the single documented requirement, an anchor element with an href attribute.

Pattern in the templateAnchor elementhref attributeGoogle can follow itTypical source
<a href="/category/boots/">Boots</a>YesYesYesHand-written template
<a href="/category/boots/"> injected by JavaScript on loadYesYesYes, the markup is what mattersClient-side rendered nav
<a onclick="goto()">Boots</a>YesNoNoAnalytics-driven nav
<span class="nav-link" data-url="…">Boots</span>NoNoNoComponent library without a link primitive
<a routerLink="/boots">Boots</a> with no rendered hrefYesNoNoSPA router defaults
<button> that pushes history stateNoNoNoFaceted filters and load-more controls
<a href="#"> with the destination in JavaScriptYesPlaceholder onlyNoLegacy dropdowns
Compliance against the one documented rule. Every "No" row is navigation a user can see and a crawler cannot.

The pattern worth noticing is that three of the four failing rows still look like links in a browser and still pass a visual QA. They fail only against markup inspection or a crawl. That is why an architecture problem so often presents as a mysterious indexing problem: the site map on the whiteboard is correct and the link graph the crawler sees is a fraction of it.

An audit that tests architecture rather than describes it

  1. Crawl the site with JavaScript rendering enabled and again with it disabled. Pages that appear only in the rendered crawl are reachable only through rendering, which is slower and not guaranteed.
  2. Export click depth from the homepage for every indexable URL. Plot the distribution. A long tail beyond depth five on a site of any size is the finding.
  3. Take the median number of distinct internal link targets per page and compare it against the depth table above. A large gap means duplicated navigation, not a large site.
  4. Grep the templates for clickable elements that are not anchors: onclick handlers, spans and divs with click bindings, buttons that navigate, and router directives that do not render an href.
  5. List every URL parameter the site emits and mark which ones change the content. Any parameter that does not change content is multiplying URLs for no gain.
  6. Check that category and hub pages link down to their members and that members link back up. One-way hierarchies strand inventory.
  7. Finish by listing every URL with no internal link pointing at it. Those pages are orphaned, and no amount of architecture work above them will reach them.

What this does not fix

Architecture decides what gets found and how often it gets revisited. It does not make a page worth ranking, and it will not lift a site out of a quality problem. If a site was hit by a spam or core update, flattening the click depth is maintenance, not recovery, and it should be sequenced after the quality work rather than in place of it. The order matters, because a faster crawl of pages Google has already judged is not an improvement.

Related reading

Leave a Reply

Your email address will not be published. Required fields are marked *