Statistics for SEO: The Practical Path, from A/B Tests to Data Traps

Anyone working in SEO and marketing lives with numbers all day long: conversions climbing, traffic collapsing, one variant that seems to beat the other. The problem is not a lack of data — there is far too much of it — but knowing which numbers are telling us something true and which are merely playing tricks. Statistics applied to SEO is not an academic exercise: it is the filter that separates defensible decisions from shamanistic practices, the method that tells us when a result is worth acting on and when it is just noise dressed up as a signal. Without that filter, every report becomes a Rorschach test in which each person reads whatever they want to see.

Putting statistics at the service of SEO, however, does not mean becoming a statistician. It means having the right tools within reach for the questions that recur in everyday work: does this change really work? How much data is needed before that can be said? Is this week’s drop in traffic an anomaly or the tail of a seasonal pattern? And why do two numbers tell one story when summed together and the opposite story when kept apart?

This page is the map of those tools, organised by problem rather than by theory. The concepts are not re-explained here: each stage is an article on the blog, and they have been grouped into the four situations an SEO professional runs into most often — experimenting to choose, reading traffic over time, recognising the traps in the data, and simulating what cannot be observed directly. Whoever has a specific problem can jump to the relevant section; whoever wants to build a method can follow them in order, because it is also the order in which one skill prepares the next. Let us start with the most requested one: deciding between two alternatives with an experiment, rather than by gut feeling.

Experimenting

The cleanest way to know whether a change works is to test it against the alternative, under controlled conditions, letting the data decide.
This section gathers the tools of the experiment: the method, the two calculators needed before and after, and the most insidious mistake people fall into when they believe they have understood everything.
A well-designed experiment is worth more than a thousand opinions in a meeting: it is the only way to establish a cause instead of merely observing a coincidence.

A/B testing is the heart of this whole block. Two variants of a page, random assignment of visitors to one or the other, a rigorous comparison of the results: it is the discipline that turns the question “which version converts better?” into something measurable and defensible. Understanding how a clean test is set up — and what ruins it — is the prerequisite for everything that follows.

Even before launching a test, though, one question must be asked: how many visitors are needed for the result to mean anything? The sample size calculator for A/B tests answers precisely this, indicating how much data to collect in order to detect a difference of a given magnitude with the desired confidence. It is the tool that prevents the most common and costly mistake: closing a test too early, on numbers too small to say anything at all.

Once a test is complete, the baton passes to the other tool. The significance calculator for A/B tests takes the collected numbers — visitors and conversions for each variant — and says whether the observed difference is statistically solid or compatible with pure chance. It is the moment when the decision is made, with a criterion rather than a feeling, about whether the winning variant really won.

There is finally one last trap, and it arrives just when we believe we have the method under control. The peeking problem describes what happens when one glances at the result of an A/B test before the end, stopping the instant the data agrees: false positives inflate in silence, even though every single calculation is impeccable. It is the stage that teaches caution toward one’s own haste — the moment chosen to look matters as much as the number that is seen.

Traffic over time

Much of SEO data consists not of snapshots but of sequences: visits day by day, impressions week after week, positions oscillating over time. Reading them as isolated numbers means losing exactly the information that matters, namely how they change.
This section gathers the tools for making sense of what moves along the axis of time: the underlying trend, the recurring seasonality, and the deviations that deserve an alarm.
A number makes sense only within its trajectory: the same value can be a triumph or a disaster depending on what preceded it.

Time series analysis is the starting point for anyone looking at a traffic chart over time. It teaches how to decompose a sequence into its ingredients — trend, seasonality, the irregular part — and, with the Holt-Winters method, how to project it forward into a forecast. It is what separates “it feels like it’s dropping” from a grounded estimate of where traffic is heading.

Once we know what to expect from a series, we can recognise when something escapes the rule. Anomaly detection is the discipline that pinpoints the out-of-place points — a suspicious spike, a sudden collapse, a value the model would never have predicted. For anyone monitoring a site it is invaluable: it turns a board of metrics into a system that signals on its own when it is worth going to see what happened, instead of noticing once the damage is done.

The traps in the data

Even with the right tools, data has refined ways of deceiving us. Not because it lies — the numbers are what they are — but because the way we aggregate, compare or interpret them can lead to conclusions exactly the reverse of reality.
This section gathers two classic traps, among the most frequent in marketing reports, and teaches how to recognise them before a wrong decision has been built on top.
Data almost never lies; we are the ones who read it wrong, and it is precisely when it looks clearest that suspecting we have understood it backwards is wisest.

Simpson’s paradox is the most spectacular of these snares: a trend that appears clear-cut in the aggregate data can reverse when it is broken down into the individual groups. A channel that seems to convert better overall can be worse in every single segment, simply because of how the volumes are distributed. It is the warning that every average must be viewed with suspicion: before concluding, it always pays to ask what the data hides beneath the aggregate.

Regression to the mean is the other trap, subtler because it disguises itself as a result. When a page has an exceptional peak, the following measurement is likely to be lower — not because anything has worsened, but by simple statistics: extreme values tend to be followed by values closer to the average. Mistaking this physiological return for the effect of one of our actions is among the most common self-deceptions of those who analyse marketing data, and recognising it avoids taking credit — or blame — that does not exist.

Simulating

Sometimes the question concerns not data already in hand, but scenarios that cannot be observed: what would happen if an uncertain situation were repeated a thousand times, which outcomes might be expected, how likely the extreme cases are. When exact mathematics becomes too tangled, chance can be made to speak in a controlled way.
Simulating means interrogating uncertainty by making it happen thousands of times on a computer, instead of trying to tame it with a closed formula that often does not exist.

The Monte Carlo method is the prime tool for this kind of problem. The idea is disarming in its simplicity: instead of computing a result in exact form, the uncertain phenomenon is simulated a great many times and the distribution of outcomes is examined. From estimating a plausible traffic interval to assessing the risk of a scenario, it is a toolbox that proves useful whenever uncertainty is too tangled to be resolved on paper.

Where to begin

For someone starting from scratch who wants a single entry point, it is A/B testing: it is the skill that gives the most immediate return on daily work, because it turns everyday choices — a headline, a call to action, a layout — into measurable decisions instead of bets. From there, the two calculators are the natural complement, one before and one after every test.

There is, however, a level beneath the entire “Experimenting” block that, sooner or later, is worth confronting: why those methods work. A/B tests, sample size and significance rest entirely on the framework of inferential statistics — hypotheses, p-values, confidence intervals — and whoever wants to stop applying the calculators as a black box finds that foundation in the path dedicated to inferential statistics, the toolbox from which this entire practical path, sooner or later, ends up drawing. This is one of the thematic paths being built to navigate the blog’s articles: not new explanations, but maps that line up what is already there, starting from the real problems of those who do SEO.