SEO Split Testing: What Google’s Rules Allow, and the Sample Size That Decides Whether Your Test Proves Anything
· Royking Niba
SEO split testing means changing something on one group of URLs, leaving a comparable group unchanged, and measuring the difference in organic performance between the two groups. It is the only way to know whether a template change earned its cost, and it has two separate failure modes. The first is breaking Google’s rules while running the test. The second, far more common, is running a test that could never have detected the effect you were looking for. Google’s documentation solves the first problem for you. Nothing but arithmetic solves the second.
Google documents this, and the rules are short
There is a Search Central page on website testing, and it is unusually direct. On serving different content to crawlers, it says “don’t show one set of URLs to Googlebot, and a different set to humans. This is called cloaking, and is against our spam policies.” On marking up alternate URLs, it prefers a canonical over a noindex, because a canonical more closely matches your intent. On sending users between variants, it says to “use a 302 (temporary) redirect, not a 301 (permanent) redirect”. And on duration, it says to “run the experiment only as long as necessary” and warns that leaving a test running unnecessarily long may be interpreted as an attempt to deceive search engines.
Two of those instructions are load-bearing and get ignored. A 301 between test variants tells Google to index the destination and forget the source, which is the opposite of what a temporary experiment wants. And splitting by user agent rather than by URL is not a clever implementation detail, it is the thing the spam policies define as cloaking: “the practice of presenting different content to users and search engines with the intent to manipulate search rankings and mislead users”. The same policy page defines sneaky redirecting as “showing search engines one type of content while redirecting users to something significantly different”. A user-agent split does both by construction.
Test mechanics against what the documentation requires
| Mechanic | How it splits | Status against Google’s documented rules |
|---|---|---|
| Split by URL, variants on their own URLs, canonical to the original | By URL | The documented approach. Canonical is preferred over noindex. |
| Server-side split of pages into two groups, each group served consistently to everyone | By URL group | Compliant, and the only shape that supports a clean measurement of ranking effects. |
| 302 redirect from original to variant for a share of users | By user, temporary | Explicitly what Google names: 302, not 301. |
| 301 redirect between variants | By user, permanent | Wrong code. Tells Google to index the destination and drop the source. |
| Serve variant A to Googlebot, variant B to humans | By user agent | Cloaking, against the spam policies. |
| Client-side JavaScript swap after load, same URL | By user | Not cloaking if both versions are equivalent, but it measures user behaviour, not ranking. It cannot test a ranking hypothesis. |
| Test left running for months after a decision was reached | Any | Against the duration guidance and the stated deception risk. |
The row people argue about is the last but one. A JavaScript swap is not automatically a violation, and it is a perfectly good way to test conversion. It just cannot answer “did this change our rankings”, because Googlebot and the user are not seeing the test in the same way, and because the ranking signal you are trying to move is attached to the URL you did not vary.
The part nobody does: checking whether the test can detect anything
An SEO split test compares two groups of pages, so its resolution is set by how many pages are in each group and how much they vary between themselves. Organic click counts across a set of pages vary enormously, and that variance is the noise your effect has to beat. Here is the worked example. It is a constructed case, not a client’s data, but the inputs are unremarkable.
Take a template used on 400 pages, split 200 into the test arm and 200 into the control arm. The pages average 120 organic clicks a month each. The standard deviation across pages is 90 clicks, which is typical for a set of category or article pages where a handful of pages carry most of the traffic. The standard error of each arm’s mean is 90 divided by the square root of 200, or 6.36 clicks. The standard error of the difference between the two arms is that multiplied by the square root of two, which is 9.00 clicks.
| Pages per arm | Standard error of the difference | Smallest detectable lift at 50% power | Smallest detectable lift at 80% power |
|---|---|---|---|
| 200 | 9.00 clicks | 17.6 clicks, a 14.7% lift | 25.2 clicks, a 21.0% lift |
| 1,000 | 4.02 clicks | 7.9 clicks, a 6.6% lift | 11.3 clicks, a 9.4% lift |
| 1,730 | 3.06 clicks | 6.0 clicks, a 5.0% lift | 8.6 clicks, a 7.1% lift |
| 3,528 | 2.14 clicks | 4.2 clicks, a 3.5% lift | 6.0 clicks, a 5.0% lift |
Read the first row and the point lands. With 200 pages per arm and that spread, the smallest lift you could distinguish from noise with any real confidence is around 21 percent. A 5 percent improvement in organic clicks, which would be a genuinely good result from a title template change, is invisible at that sample size. To detect 5 percent with 80 percent power you need about 3,528 pages in each arm, which is 7,056 pages on that one template. Most sites running SEO split tests do not have that, and the honest conclusion is that their test can only ever confirm large effects.
What that arithmetic does not cover
This is a simple two-sample comparison and it deliberately ignores three things that matter in practice. Seasonality moves both arms together, which is why the comparison is between arms rather than against last month. Pages are not independent, because a template change can move internal linking and cannibalisation across the whole set. And click counts are heavily skewed rather than normally distributed, so in a real test you would work on a transformed metric or use each page’s own pre-period as its baseline, which cuts the variance considerably and is the main reason production SEO testing platforms beat a naive split. Those methods lower the numbers in that table. They do not change the shape of the problem, and none of them rescue 200 pages per arm looking for 5 percent.
How I would set one up
- Write the hypothesis and the effect size you would act on before anything else. If a 3 percent lift would not change your decision, do not design a test that only resolves 3 percent.
- Count the pages on the template and run the arithmetic above. If the smallest detectable lift is larger than the effect you expect, stop. You are choosing between a bigger sample, a bolder change, and not testing.
- Split by URL group, server-side, and assign randomly but stratify by traffic band. Random assignment across a skewed set can easily put most of your big pages in one arm.
- Keep everything else frozen for the duration. A migration, a redesign or a core update landing mid-test ends the test rather than complicating it.
- Serve every visitor and every crawler the same thing for a given URL. If you must redirect users, use a 302.
- Measure in Search Console by page group, not by sampled analytics sessions. Clicks and impressions per URL are the units the hypothesis is about.
- End the test when it is decided, then remove the test apparatus. Google’s duration guidance is not a formality, and an experiment left up for a year looks like something else.
The uncomfortable conclusion from all of this is that most sites cannot run a meaningful SEO split test on most of their templates, and that is worth knowing before spending a quarter on one. The sites that can are the ones with thousands of pages on a single template. For everyone else, the honest alternatives are a staged rollout with a pre-period baseline, or accepting that some changes are made on documentation and judgement rather than on measurement.
Related reading
- 301 vs 302 redirects: which one Google canonicalizes, which is the mechanism the testing guidance depends on.
- Cloaking in SEO: what counts and what does not, for the line a user-agent split crosses.
- Canonical tag SEO and why Google ignores yours, because a canonical on a test variant is a hint rather than an instruction.
- Google penalty recovery: how I diagnose a traffic collapse, for when a test is suspected of having caused one.
- The SEO audit checklist I actually use, which is where a testing programme should start rather than finish.
Leave a Reply