Statistical significance is a way of judging whether a difference observed in a test — one landing page converting better than another, one creative outperforming another — is large enough, relative to the amount of data and its variability, that it is unlikely to be just random chance.
Definition
Statistical significance is a way of judging whether a difference observed in a test — one landing page converting better than another, one creative outperforming another — is large enough, relative to the amount of data and its variability, that it is unlikely to be just random chance. It is usually expressed through a p-value (the probability of seeing a difference this big if there were really no difference) or a confidence level, with a threshold agreed in advance (commonly 95%).
Significance answers only one narrow question: is the effect probably real. It does not say the effect is large, important, or stable over time, and it depends heavily on sample size — with enough traffic a trivial difference becomes significant, and with too little traffic a real difference stays undetectable.
It is a gate against acting on noise, not a measure of business value.
In context
For iGaming affiliates and operators running split tests, significance discipline is what stops teams from shipping changes based on early, lucky-looking results. The common mistakes: stopping a test the moment it crosses 95% (peeking inflates false positives), running many variants at once without adjusting the threshold, and judging on an upstream metric (registration rate) that reaches significance quickly while the metric that matters (cost per FTD, retention) needs far more data and time because deposits arrive days after the click.
The practical approach is to decide the primary metric, the minimum effect worth detecting, and the required sample size before starting; run the test to that sample without peeking to decide; and then look at whether the result is both significant and big enough to act on. In this sector the deposit lag means many tests need to run for weeks, and a page that wins significantly on sign-ups can still lose on player value, so significance on the right, downstream metric is what counts.
Significance also is not permanence — seasonality, traffic-source shifts and audience changes can reverse a result — so important wins are worth re-validating. Used properly, significance testing keeps optimisation honest; used as a green light to grab at the first favourable p-value, it manufactures confident decisions from randomness.
Worked example
An affiliate split-tests two prelanders and sees the new one hit 96% significance on registration rate after three days. Instead of shipping, the team continues to the pre-agreed sample and the 14-day deposit window; by then the new prelander is significant on registrations but slightly worse on cost per FTD, so it is not adopted.
Related terms
Frequently asked questions
Browse the full iGaming & affiliate glossary — hundreds of EN/RU terms with examples.
← Back to glossary