Royking Niba

Orphan Pages: Why Google Never Finds Them, and the Audit That Surfaces Yours

· Royking Niba

Stock photograph of a laptop on a desk showing an analytics dashboard. It is not a screenshot of any client's real reporting.

An orphan page is a URL on your site that no other page on your site links to. It exists, it may be published, it may even be indexed, and nothing in your own navigation or body copy points at it. The reason this matters is not a ranking penalty, because there is no such thing as an orphan page penalty. It matters because of how Google says it finds URLs in the first place.

What Google’s documentation actually says about discovery

The How Search Works documentation describes URL discovery in two sentences that are worth reading literally. Google says that “some pages are known because Google has already visited them” and that “other pages are discovered when Google extracts a link from a known page to a new page”, along with the case where “you submit a list of pages (a sitemap) for Google to crawl”. That is the whole list. A page with no internal link, no external link and no sitemap entry has no documented route into Google’s crawl queue at all.

The sitemaps documentation then removes the comfort most teams take from having a sitemap. Google states that “a sitemap helps search engines discover URLs on your site, but it doesn’t guarantee that all the items in your sitemap will be crawled and indexed”. It also tells you when a sitemap is the thing that saves you: on large sites, Google notes, “it’s more difficult to make sure that every page is linked by at least one other page”, and on new sites “Googlebot might not discover your pages if no other sites link to them”. Read the other way round, the same page says you may not need a sitemap at all if your site is small and “comprehensively linked internally”, because “Googlebot can find all the important pages on your site by following links”.

So the documentation treats internal linking as the primary discovery mechanism and the sitemap as a supplement that carries no guarantee. An orphan page is a page that has opted out of the primary mechanism and is relying entirely on the one with the disclaimer attached.

Discovery sources, and what each one actually buys you

Teams conflate three different things: whether Google can find a URL, whether Google will prioritise crawling it, and whether the URL receives any internal ranking signal. A sitemap entry does the first and nothing else. This is the matrix I use in audits.

Discovery sourceGets the URL foundPasses internal signalSignals importance
Internal link from a crawlable pageYesYesYes, through link position and count
XML sitemap entryYes, with no guarantee of crawling or indexingNoNo
External link from another siteYesYes, from outsideYes
Redirect pointing at the URLYesYes, from the source URLPartly
Canonical tag pointing at the URLYesConsolidates from the duplicateNo, it is a hint
Nothing at allNo documented routeNoNo

The row that catches people is the second one. A page that appears in your sitemap and nowhere else is discoverable and completely unsupported. It has no internal links, so it has no internal signal, and Google’s own crawl-prioritisation behaviour has nothing to work with beyond the fact that you listed it.

The arithmetic, on a site shaped like the ones I audit

The numbers below are a worked example rather than a client’s data, but the shape is the one that keeps turning up on mid-sized publishers and stores. Work through it and the size of the problem becomes obvious in a way that a crawl report full of warnings never manages.

MeasurementURLsShare of published
Published, indexable URLs in the CMS12,000100%
Reachable by following internal links from the homepage8,40070.0%
Listed in the XML sitemap11,20093.3%
In the sitemap but not reachable by internal links2,80023.3%
In neither the link graph nor the sitemap8006.7%
Distinct URLs Googlebot requested in 30 days of logs9,10075.8%

Three findings fall straight out of those six rows. First, 3,600 URLs are orphaned by the link graph, which is 30 percent of the site. Second, 800 of those have no documented discovery route whatsoever, so they are not orphans so much as invisible. Third, and this is the number that usually changes the conversation, Googlebot requested 9,100 distinct URLs while only 8,400 are link-reachable, so at most 700 of the 2,800 sitemap-only URLs were fetched in a month. Around 2,900 published URLs went unrequested.

Notice what the log line does that a crawler cannot. A site crawl tells you which URLs are reachable. Search Console tells you what Google concluded about the URLs it processed. Only the server log tells you which URLs Googlebot actually asked for, which is the only way to separate “orphaned and ignored” from “orphaned but found anyway”.

Where orphan pages come from

  • Pagination that was canonicalised away. If every page of a paginated sequence canonicalises to page one, the listing links on pages two and beyond stop being the route to the items they list.
  • Faceted URLs generated by filters. These are usually the opposite problem, a flood of crawlable URLs, but the products that only ever appeared under a now-blocked filter combination lose their only link.
  • Imported or migrated content. A bulk import creates posts with no category, no tag and no listing, and the sitemap is the only thing that ever mentions them.
  • Landing pages built for paid or email campaigns. Deliberately unlinked, which is fine, until someone expects them to rank.
  • Retired navigation. A menu rebuild drops a section, the URLs stay published, and nothing links to them the next day.
  • Author, date and tag archives set to noindex. Noindexing an archive is often correct, and it also removes it as a discovery route for everything it listed.

The last two are the ones that produce orphans quietly, because nothing about them looks like a mistake at the time.

The audit, in the order I run it

  1. Get the true published list from the database or the CMS, not from a crawl. A crawl can only ever return what it can reach, which is the exact thing you are trying to measure.
  2. Crawl the internal link graph from the homepage with sitemaps switched off. That set is your link-reachable population.
  3. Subtract. Published minus link-reachable is your orphan set. Do this before looking at any tool’s own orphan report, so you own the number.
  4. Split the orphan set by discovery route. Sitemap-only, external-link-only, and nothing at all. The three groups need different fixes.
  5. Check the logs for 30 days. For each orphan, was it requested by a verified Googlebot at all. This is where the set shrinks to the ones that genuinely are not being reached.
  6. Triage on value, not on tidiness. An orphan that should not exist gets removed or consolidated. An orphan that should rank gets a real internal link from a page that is itself linked and crawled.
  7. Fix the generator, not the instances. If the cause is a template, a canonical rule or a retired menu, linking the current orphans by hand guarantees you are back here next quarter.

What an orphan page is not

Two clarifications, because both come up in every engagement. There is no orphan page penalty and Google has never published one, so a client whose traffic fell is not being demoted for having orphans. And a sitemap-only URL is not automatically a problem: if the page is intentionally unlinked, such as a paid landing page, then it is doing exactly what you asked. The problem is only a problem when a page you need to rank has been cut off from the mechanism Google documents as its primary route to finding pages. That is worth repeating because Google is equally explicit that discovery is not the finish line: “Google doesn’t guarantee that it will crawl, index, or serve your page, even if your page follows the Google Search Essentials.”

Related reading

Leave a Reply

Your email address will not be published. Required fields are marked *