Royking Niba

International SEO: The URL Structure Decision, and the hreflang Arithmetic Behind It

· Royking Niba

Stock photograph of a desk globe in greyscale. It is a licensed stock image and not a map of any particular site or market.

International SEO looks like a large problem and is really two decisions. The first is where the localised versions of a page live. The second is whether the annotations that tie those versions together are reciprocal, because Google ignores the ones that are not. The first decision is strategic and is made once. The second is arithmetic, it scales badly, and it is where most multilingual sites quietly lose the benefit of the work.

The four URL structures, and what Google documents about each

Google’s multi-regional and multilingual sites guide lists four ways to split a site by country or language, with the trade-offs for each. Three of them are viable. One is dismissed in two words.

StructureExampleDocumented strengthsDocumented weaknesses
Country-specific domainexample.deClear geotargeting, server location irrelevantExpensive, requires more infrastructure
Subdomain on a generic domainde.example.comEasy to set up, allows different server locationsUsers might not recognise geotargeting from the URL alone
Subdirectory on a generic domainexample.com/de/Easy to set up, low maintenance on the same hostSingle server location, separation of sites is harder
URL parameterexample.com?loc=deNone listedNot recommended, URL-based segmentation is difficult
Compiled from Google’s Managing multi-regional and multilingual sites documentation, checked 1 October 2026.

Three things that are commonly treated as levers are not levers at all. Google states that for working out the language of a page "we don’t use any code-level language information such as lang attributes, or the URL", and relies instead on the "visible content of your page". It also ignores locational meta tags. And on the single most common piece of international engineering, the language sniffer, the guidance is a plain instruction.

Avoid automatically redirecting users from one language version of a site to a different language version.

Google Search Central, Managing multi-regional and multilingual sites

A redirect based on detected language takes the decision away from the user and, more to the point here, takes it away from the crawler too. If every request from a United States address is bounced to the English version, the German version has no reliable route to being fetched at all.

hreflang is a reciprocal system, and that is where it fails

The annotations themselves are simple. Each version carries a code made of "the language code (in ISO 639-1 format) followed by an optional second code that represents the region code (in ISO 3166-1 Alpha 2 format)", there is a reserved x-default value "used when no other language/region matches the user’s browser setting", and "each language version must list itself as well as all other language versions". The rule that does the damage is the next one.

If two pages don’t both point to each other, the tags will be ignored.

Google Search Central, Tell Google about localized versions of your page

That is not a warning about quality, it is a statement that the relationship is discarded. A one-way annotation is worth exactly nothing, and the documentation adds that where the requirement is not met for every page using annotations, "those annotations may be ignored or not interpreted correctly". So the useful question is not whether your hreflang is correct. It is how many reciprocal pairs your site is carrying, and what a small error rate across them actually costs.

The arithmetic of eight locales

Take a mid-sized site: eight locales and two thousand page sets, meaning two thousand pieces of content that exist in all eight. Nothing about that is unusual for a European business. Here is what it generates.

QuantityValueHow it is derived
Locales8Given
Translated page sets2,000Given
Total pages16,0008 locales x 2,000 sets
Annotations per page98 versions including the page itself, plus x-default
Total annotation lines144,00016,000 pages x 9
Reciprocal pairs per set288 x 7 divided by 2
Reciprocal pairs sitewide56,00028 pairs x 2,000 sets
Broken lines at a 1% error rate1,4401% of 144,000
Page sets affectedBetween 52 and 1,4401,440 divided by 28 if the errors cluster, one per set if they spread
Worked example. Assumes every page set exists in all eight locales and that each broken line destroys one reciprocal pair.

The last row is the finding. A one percent error rate, which almost any engineering team would sign off as excellent, affects somewhere between 2.6 and 72 percent of your content depending on nothing more than whether the faults cluster in one locale or scatter across the catalogue. You cannot tell which from the error rate alone, and nothing in Search Console reports it that way. You have to go and count.

The reason the spread is so wide is combinatorial. Annotations grow with the number of locales, but reciprocal pairs grow with its square. Going from four locales to eight doubles the pages and doubles the lines, and it quadruples the pairs that must all hold: six per set becomes twenty-eight. That is the real cost of adding a market, and it is never the cost that appears in the business case.

Where to put the annotations

Google documents three delivery methods and treats them as equivalent: "there are three ways to indicate multiple language/locale versions of a page to Google: HTML, HTTP Headers, Sitemap". They are equivalent to the crawler and nothing like equivalent to maintain.

MethodWhere it livesAdded to each pageWhat it means at 16,000 pages
HTML link elementsThe head of every page9 elements144,000 elements spread across every template, changed on every market launch
HTTP headersThe response for each URL9 header valuesThe only option for non-HTML files such as PDFs
SitemapOne or more sitemap filesNothingThe whole set lives in one generated artefact, which is why large sites choose it
Method list from Google’s localized versions documentation. The scale column applies the eight-locale worked example above.

The sitemap method wins on sites of this size for a reason that has nothing to do with crawling and everything to do with failure modes. Annotations generated in one place from one source of truth either break everywhere at once, which you will notice, or not at all. Annotations embedded in templates break one locale at a time, silently, usually on the launch of the ninth market.

The audit I run on an international site

  1. Count the reciprocal pairs, not the annotations. Crawl every locale, build the directed graph, and list every pair where only one direction exists.
  2. Check that every page lists itself. A missing self-reference is the most common single fault and it is invisible in a spot check.
  3. Validate the codes. Language in ISO 639-1, optional region in ISO 3166-1 Alpha 2. The usual casualty is a region code invented from a language name.
  4. Confirm exactly one x-default across the set, pointing at the version you actually want served when nothing matches.
  5. Check that annotated URLs are indexable and canonical to themselves. An annotation pointing at a page that canonicalises elsewhere asks Google to do two contradictory things.
  6. Test for automatic language redirects from several regions. If a locale cannot be reached without a redirect, its annotations cannot be verified either.
  7. Pick one delivery method and remove the others. Sites that acquire a second method during a migration end up with two sets of annotations that disagree, and the disagreement is what gets ignored.

None of this is exotic. It is counting, done properly, on a system where the penalty for a missing edge is silent. That is the whole of international SEO once the URL structure decision is behind you.

Related reading

Leave a Reply

Your email address will not be published. Required fields are marked *