Royking Niba

SEO Split Testing: What Google’s Rules Allow, and the Sample Size That Decides Whether Your Test Proves Anything

· Royking Niba

Stock photograph of a walking trail in a forest forking into a narrow path and a wide one.

SEO split testing means changing something on one group of URLs, leaving a comparable group unchanged, and measuring the difference in organic performance between the two groups. It is the only way to know whether a template change earned its cost, and it has two separate failure modes. The first is breaking Google’s rules while running the test. The second, far more common, is running a test that could never have detected the effect you were looking for. Google’s documentation solves the first problem for you. Nothing but arithmetic solves the second.

Google documents this, and the rules are short

There is a Search Central page on website testing, and it is unusually direct. On serving different content to crawlers, it says “don’t show one set of URLs to Googlebot, and a different set to humans. This is called cloaking, and is against our spam policies.” On marking up alternate URLs, it prefers a canonical over a noindex, because a canonical more closely matches your intent. On sending users between variants, it says to “use a 302 (temporary) redirect, not a 301 (permanent) redirect”. And on duration, it says to “run the experiment only as long as necessary” and warns that leaving a test running unnecessarily long may be interpreted as an attempt to deceive search engines.

Two of those instructions are load-bearing and get ignored. A 301 between test variants tells Google to index the destination and forget the source, which is the opposite of what a temporary experiment wants. And splitting by user agent rather than by URL is not a clever implementation detail, it is the thing the spam policies define as cloaking: “the practice of presenting different content to users and search engines with the intent to manipulate search rankings and mislead users”. The same policy page defines sneaky redirecting as “showing search engines one type of content while redirecting users to something significantly different”. A user-agent split does both by construction.

Test mechanics against what the documentation requires

MechanicHow it splitsStatus against Google’s documented rules
Split by URL, variants on their own URLs, canonical to the originalBy URLThe documented approach. Canonical is preferred over noindex.
Server-side split of pages into two groups, each group served consistently to everyoneBy URL groupCompliant, and the only shape that supports a clean measurement of ranking effects.
302 redirect from original to variant for a share of usersBy user, temporaryExplicitly what Google names: 302, not 301.
301 redirect between variantsBy user, permanentWrong code. Tells Google to index the destination and drop the source.
Serve variant A to Googlebot, variant B to humansBy user agentCloaking, against the spam policies.
Client-side JavaScript swap after load, same URLBy userNot cloaking if both versions are equivalent, but it measures user behaviour, not ranking. It cannot test a ranking hypothesis.
Test left running for months after a decision was reachedAnyAgainst the duration guidance and the stated deception risk.

The row people argue about is the last but one. A JavaScript swap is not automatically a violation, and it is a perfectly good way to test conversion. It just cannot answer “did this change our rankings”, because Googlebot and the user are not seeing the test in the same way, and because the ranking signal you are trying to move is attached to the URL you did not vary.

The part nobody does: checking whether the test can detect anything

An SEO split test compares two groups of pages, so its resolution is set by how many pages are in each group and how much they vary between themselves. Organic click counts across a set of pages vary enormously, and that variance is the noise your effect has to beat. Here is the worked example. It is a constructed case, not a client’s data, but the inputs are unremarkable.

Take a template used on 400 pages, split 200 into the test arm and 200 into the control arm. The pages average 120 organic clicks a month each. The standard deviation across pages is 90 clicks, which is typical for a set of category or article pages where a handful of pages carry most of the traffic. The standard error of each arm’s mean is 90 divided by the square root of 200, or 6.36 clicks. The standard error of the difference between the two arms is that multiplied by the square root of two, which is 9.00 clicks.

Pages per armStandard error of the differenceSmallest detectable lift at 50% powerSmallest detectable lift at 80% power
2009.00 clicks17.6 clicks, a 14.7% lift25.2 clicks, a 21.0% lift
1,0004.02 clicks7.9 clicks, a 6.6% lift11.3 clicks, a 9.4% lift
1,7303.06 clicks6.0 clicks, a 5.0% lift8.6 clicks, a 7.1% lift
3,5282.14 clicks4.2 clicks, a 3.5% lift6.0 clicks, a 5.0% lift

Read the first row and the point lands. With 200 pages per arm and that spread, the smallest lift you could distinguish from noise with any real confidence is around 21 percent. A 5 percent improvement in organic clicks, which would be a genuinely good result from a title template change, is invisible at that sample size. To detect 5 percent with 80 percent power you need about 3,528 pages in each arm, which is 7,056 pages on that one template. Most sites running SEO split tests do not have that, and the honest conclusion is that their test can only ever confirm large effects.

What that arithmetic does not cover

This is a simple two-sample comparison and it deliberately ignores three things that matter in practice. Seasonality moves both arms together, which is why the comparison is between arms rather than against last month. Pages are not independent, because a template change can move internal linking and cannibalisation across the whole set. And click counts are heavily skewed rather than normally distributed, so in a real test you would work on a transformed metric or use each page’s own pre-period as its baseline, which cuts the variance considerably and is the main reason production SEO testing platforms beat a naive split. Those methods lower the numbers in that table. They do not change the shape of the problem, and none of them rescue 200 pages per arm looking for 5 percent.

How I would set one up

  1. Write the hypothesis and the effect size you would act on before anything else. If a 3 percent lift would not change your decision, do not design a test that only resolves 3 percent.
  2. Count the pages on the template and run the arithmetic above. If the smallest detectable lift is larger than the effect you expect, stop. You are choosing between a bigger sample, a bolder change, and not testing.
  3. Split by URL group, server-side, and assign randomly but stratify by traffic band. Random assignment across a skewed set can easily put most of your big pages in one arm.
  4. Keep everything else frozen for the duration. A migration, a redesign or a core update landing mid-test ends the test rather than complicating it.
  5. Serve every visitor and every crawler the same thing for a given URL. If you must redirect users, use a 302.
  6. Measure in Search Console by page group, not by sampled analytics sessions. Clicks and impressions per URL are the units the hypothesis is about.
  7. End the test when it is decided, then remove the test apparatus. Google’s duration guidance is not a formality, and an experiment left up for a year looks like something else.

The uncomfortable conclusion from all of this is that most sites cannot run a meaningful SEO split test on most of their templates, and that is worth knowing before spending a quarter on one. The sites that can are the ones with thousands of pages on a single template. For everyone else, the honest alternatives are a staged rollout with a pre-period baseline, or accepting that some changes are made on documentation and judgement rather than on measurement.

Related reading

Leave a Reply

Your email address will not be published. Required fields are marked *