analytics
How Long Should an A/B Test Run?
Two identical popups can still post different daily numbers by pure chance. Here's the statistics behind how long a test really needs to run, with charts.
A test that shows a gap on day three usually isn't showing you a winner — it's showing you randomness. Two genuinely identical popups will still produce different daily numbers, and how big a real difference actually is determines how much traffic it takes to tell it apart from that noise. Here's the intuition behind why, with the charts to go with it.
Two identical popups still look different, every single day
Imagine you split your traffic in half and show both halves the exact same popup — same discount, same everything. Nothing at all is different between Group A and Group B. If you plot each group's daily accept rate, you'd expect two flat, identical lines.
That's not what happens.
Illustrative example — both lines are drawn from the exact same true rate. Every wobble is ordinary day-to-day variation, not a real difference between the popups.
This isn't a flaw in how a test works. It's what randomness looks like. In a simulation like this one, a visible gap between two truly identical groups shows up on any given day roughly three times out of ten — not because one popup is better, but because a coin doesn't land exactly 50/50 every ten flips either. A gap on day one, or day three, or even day five, tells you nothing on its own.
The coin-flip intuition: obvious biases are easy, small ones aren't
Here's the same idea in its simplest form.
Imagine a coin that lands heads 80% of the time. Flip it ten times and you'll see it — seven, eight, maybe nine heads. Nobody needs statistics to notice that.
Now imagine a coin that lands heads 51% of the time — barely biased at all. Flip it ten times and you'll get somewhere between three and eight heads, completely indistinguishable from a fair coin. You'd need to flip it thousands of times before that tiny lean toward heads reliably shows up above the noise.
Your popup test works the same way. If one version genuinely converts a lot better, it's the 80% coin — it shows up fast. If the real difference is small, it's the 51% coin: real, maybe even worth having, but it hides inside normal randomness until you've shown your two versions to a lot of shoppers. Plug your own cart volume into our A/B test duration calculator to see roughly how much traffic that takes for your store.
What your test can actually detect shrinks as traffic comes in
Every A/B test has an invisible boundary: the smallest gap between two versions it's currently capable of telling apart from noise. Early on, with only a handful of visits in each group, that boundary is wide — almost any difference you see could just be chance. As more shoppers come through, the boundary narrows, and smaller real differences become possible to confirm.
Illustrative example. The shaded zone is where a real difference and pure chance still look the same — anything inside it hasn't been confirmed yet. A big real difference pokes outside the zone almost immediately. A small real difference stays hidden until much more traffic has accumulated.
The wait isn't a countdown — it's a shrinking margin of error
It helps to stop thinking of an A/B test as something that runs for a fixed number of days and then reports back. It's closer to a margin of error that keeps shrinking as more shoppers see your two versions.
A big real difference doesn't wait for a schedule — it confirms itself in days, sometimes faster, because it's large enough to clear even a wide early margin.
If a gap between your two versions stays too small to confirm even as your traffic keeps growing, that isn't an inconclusive test. That's the answer. It means any real difference is small enough that it doesn't matter economically. "No result" isn't a shrug — it's a finding. It means the two versions perform about the same, so keep whichever one is cheaper to run, easier to maintain, or gives away a smaller discount.
If you haven't run a structured test yet, start with the test most stores get the clearest signal from. And if your store's traffic is on the lower side, read why most Shopify A/B tests never finish before you commit to one — it walks through estimating your own realistic timeline in advance, rather than finding out three weeks in.
The three metrics we check, in order, and why the order matters
Once a test has run long enough to say something, we don't look at one number — we look at three, in a specific order, because each one answers a different question and becomes trustworthy at a different point in the test.
1. Accept rate — the earliest signal. Accept rate, how many shoppers who see the popup take the offer, is checked first because it happens for nearly everyone who sees the popup. It accumulates data fastest, so it clears the noise floor earliest.
2. Prevention rate — the metric that decides the winner. Prevention rate, orders per popup shown, decides which version actually wins. Accepting an offer isn't the same as completing a purchase, so this is the number that reflects what shoppers actually did rather than what they clicked.
3. Net revenue per intervention — the metric that can veto a win, never award one. Each time the popup is shown counts as one intervention. A version with a bigger discount can post a higher prevention rate while giving away more than the extra orders bring in. Net revenue per intervention catches that. It never crowns a winner on its own, because order values vary so much from shopper to shopper that revenue needs far more data than the other two metrics before it's trustworthy.
Check them in that order: let accept rate tell you a test is worth watching, let prevention rate tell you who won, and let net revenue per intervention have the final word on whether the win is worth having.
Frequently asked questions
-
Why does a bigger change need fewer days, and a smaller change more? It feels backwards. The test is telling a real difference apart from random day-to-day noise. A big lift stands out from the noise quickly; a small lift hides inside it, so you need many more visitors to average the noise down. Halving the lift you want to catch roughly quadruples the visitors needed. Example: catching a 20% lift on a 3% rate takes about a quarter of the traffic needed for a 10% lift.
-
What if I plan for a small change, wait the extra days, and the change turns out to be big? Nothing is lost. A big lift will usually show as significant early, and the planned duration is the maximum you need for the smallest lift worth acting on, not a wait you must serve. Keep running at least two full weeks so weekday and weekend behaviour are both covered, then read the result. The extra days only make the estimate tighter (a narrower range around the true lift).
-
Can I stop as soon as the result looks significant? Not on a single early peek. Early results swing a lot; a lift that is significant on day 3 can vanish by day 10. Decide the duration up front, run the full business cycle, then decide. If you must check earlier, treat it as a preview, not a verdict.
-
Relative vs absolute lift, what is the difference? Absolute is the change in percentage points (3% → 3.6% is +0.6 points); relative is that change divided by where you started (0.6 ÷ 3 = 20%). The calculator uses relative lift because it compares fairly across stores with different starting rates.
Start your test with the right expectations
None of this changes what you should do — run one test at a time, change one setting between your two versions, and let it run. It changes what you should expect while it's running. A clean gap on day two proves less than it looks like. A test that goes quiet for two weeks might be about to confirm the biggest result you've seen. And a test that never separates the two versions is telling you exactly as much as one that does.