Royking Niba

JavaScript SEO: The Three Phases Googlebot Uses, and the Discovery Cost of Rendering Your Links

· Royking Niba

Stock photograph of coloured source code on a dark screen. It is a licensed stock image and not a screenshot of any site discussed in this article.

JavaScript SEO is not a question of whether Google can run JavaScript. It can, and has been able to for years. The question is when it runs it, because Google does not crawl and render in a single pass. The gap between those two steps is where almost every real problem on a JavaScript site lives, and the cost of that gap compounds with every click of depth you put behind it.

What Google documents about the process

The documentation is unusually direct about the shape of the pipeline. It says plainly that "Google processes JavaScript web apps in three main phases: 1. Crawling 2. Rendering 3. Indexing", and that rendering is a queue rather than something that happens during the fetch.

Googlebot queues all pages with a 200 HTTP status code for rendering, unless a robots meta tag or header tells Google not to index the page.

Google Search Central, JavaScript SEO basics

Two consequences follow from that one sentence, and they are the two that teams miss. The first is that anything your scripts produce is invisible until the queue reaches your page. The second is that the noindex check happens before rendering, not after, which is why Google states that "when Google encounters the noindex tag, it may skip rendering and JavaScript execution, which means using JavaScript to change or remove the robots meta tag from noindex may not work".

The three phases, and what each one cannot do

PhaseWhat Googlebot doesWhat it cannot do at this point
CrawlingFetches the URL and reads the raw HTML responseSee anything a script has not yet produced
RenderingQueues the page, then runs the JavaScript in a headless browserRender a page it was told not to index, or a resource blocked in robots.txt
IndexingReads the rendered DOM and stores what it findsIndex a URL that nothing ever linked to in a crawlable way
Compiled from Google’s JavaScript SEO basics and Fix search-related JavaScript problems documentation, checked 1 October 2026.

The third row is the one that costs traffic. Discovery happens in the crawling phase, from the raw HTML, unless your links only exist after rendering. Google is explicit that "Google can only discover your links if they are <a> HTML elements with an href attribute". If those elements are injected by a script, the link still works, but it is discovered one whole pipeline stage later.

The discovery cost of putting your links behind rendering

Here is the arithmetic that makes the problem concrete. Take a site whose templates carry 24 internal links each, and assume uniform branching so that every page at a given depth links to 24 new URLs. This is a simplification, because real link graphs overlap heavily, and it is stated here rather than hidden. It is still enough to show the shape.

Click depthURLs reachable at that depthLinks in the raw HTMLLinks injected by JavaScript
124Discovered on the first crawlDiscovered after one render
2576Discovered on the second crawl roundDiscovered after two renders
313,824Discovered on the third crawl roundDiscovered after three renders
4331,776Discovered on the fourth crawl roundDiscovered after four renders
Uniform branching model at 24 crawlable links per page. Real link graphs overlap, so these are ceilings on reach, not counts of distinct URLs.

The two right-hand columns are the whole argument. Server-rendered links need one queue to clear per level. Script-injected links need two, and the second one is the queue Google does not publish a service level for. On a four-level catalogue that is the difference between discovery driven by crawl capacity and discovery driven by render capacity, and only one of those is something you can influence from your own server.

Note what this is not. It is not an argument that client-side rendering cannot rank, and Google does not say that either. It is an argument about latency and about the order in which your URLs become known, which matters most on large catalogues, on anything with a news cycle, and on any site recovering from a drop where the recrawl speed is the whole game.

Five patterns that break, and what the documentation says about each

PatternWhat Google documentsNet effect
A link written as a span or div with an onclick handler"Google can only discover your links if they are <a> HTML elements with an href attribute"The destination is never discovered from this page, rendered or not
Routing between views with a URL fragment instead of the History APIGoogle tells developers to "use the History API to implement routing between different views"Views share one URL, so only one of them can be indexed
Script or CSS files disallowed in robots.txt"Google Search won’t render JavaScript from blocked files or on blocked pages"The page is indexed as its unrendered shell
Using JavaScript to remove a noindex tag"When Google encounters the noindex tag, it may skip rendering and JavaScript execution"The page stays out of the index, and the fix never runs
A single page app returning 200 for a route that does not existSPAs "often report a 200 HTTP status code instead of the appropriate status code. This can lead to error pages being indexed and possibly shown in search results"Soft 404s accumulate in the index
Quotations from Google’s JavaScript SEO basics and Fix search-related JavaScript problems pages, checked 1 October 2026.

Three of those five still look like working links or working pages in a browser, which is why they survive QA. The browser follows an onclick handler happily. A fragment route renders the right view. A SPA error page looks like an error page to a human. In all three cases the thing that fails is invisible unless you look at what the crawler receives rather than at what the user sees.

Canonicals deserve one line of their own. Google allows setting them in script but draws a hard line on what you may change: "you shouldn’t use JavaScript to change the canonical URL to something else than the URL you specified". A canonical rewritten client-side is one of the harder faults to find, because the raw HTML and the rendered DOM disagree and most crawling tools only show you one of them.

The audit I run on a JavaScript site

  1. Diff the raw HTML against the rendered DOM for one URL per template. Everything that appears only on the right-hand side is content that waits on the render queue.
  2. Count crawlable links in the raw response. Not links in the rendered page: anchor elements with an href attribute in the bytes the server sent.
  3. Check robots.txt against your bundle paths. A Disallow on a build directory is the most common cause of a site that indexes as an empty shell.
  4. Request a route that does not exist and read the status code, not the page. A 200 on a missing route is a soft 404 factory.
  5. Compare the canonical in the source with the canonical in the DOM. They should be identical. If they are not, you have found the bug before the client has.
  6. Use the URL Inspection tool on live pages, where you "can see loaded resources, JavaScript console output and exceptions, rendered DOM, and more information". A console exception during render is a silent indexing failure.
  7. Decide what has to survive without JavaScript. In practice that list is short: the canonical, the robots meta tag, the title, the main content and the internal links. Everything else can wait for the queue.

That last step is the whole discipline compressed into one sentence. You do not have to abandon client-side rendering to fix JavaScript SEO. You have to decide, deliberately, which five things ship in the first response, and then verify that they did.

Related reading

Leave a Reply

Your email address will not be published. Required fields are marked *