NavonaAI logoNavonaAI

analytics

Why Most Shopify A/B Tests Never Finish

On a typical low-traffic store, detecting a real 10% change in your popup can take over a year. Here's the math, and how to know before you commit to a test.

NavonaAI Team4 min read

On a store doing around 33 carts a day with a 20% cart-to-order rate, detecting a real 10% relative change in a cart abandonment popup takes roughly 13 months.

Most merchants don't run their tests for 13 months. They run them for two weeks, glance at the numbers, and declare a winner. What they're usually reading is noise, not a result — and the math behind why is worth understanding before you set up your next test.

What "minimum detectable effect" actually means

Every A/B test has a smallest change it can reliably tell apart from chance. That's the minimum detectable effect, and it's a function of two things: how many carts you're testing on, and how long you're willing to wait.

A coin that's only slightly weighted toward heads still looks like an even coin after a handful of flips. Flip it enough times and the weighting becomes obvious. Cart conversion works the same way. A real 10% improvement in your popup is a small effect sitting inside a lot of day-to-day variation — different traffic sources, different days of the week, different shoppers with different intent. Separating a small real effect from that noise takes volume. Without enough of it, a "winning" variant after two weeks is often just whichever version happened to catch a better mix of shoppers.

Why low-traffic stores need bigger differences, not just more patience

This is the part that surprises people: it's not just that low-traffic stores need to wait longer. It's that the size of the difference they can detect at all shrinks as traffic shrinks.

A store with high daily cart volume can detect a modest, realistic improvement — the kind an actual optimization is likely to produce — in a reasonable window. A store with 33 carts a day is stuck choosing between two bad options: wait over a year to detect a normal-sized improvement, or only run tests where you expect a huge, obvious difference between versions. Most real optimizations aren't huge and obvious. They're incremental. Which means most low-traffic stores are set up to fail the exact tests they're most likely to want to run.

Why a test that outlives your storefront is worthless

Thirteen months is longer than most stores go without touching their site. A new collection launches. A pricing change ships. The homepage gets redesigned. Somewhere in that window, the storefront your test started on stops being the storefront your test is still measuring, and whatever result eventually arrives is answering a question about a version of your store that no longer exists.

A test only means something if it can finish while the thing it's testing stays the same. If your realistic test duration is measured in months longer than your release cycle, the test isn't going to give you a usable answer — it's going to give you a number to report on a page that's already changed underneath it.

Know this before you start, not after

The fix isn't to run tests forever or to give up on testing. It's to check, before committing to a test, whether your traffic and the effect size you're hoping to detect can actually produce a conclusion in a timeframe you'd act on. If the honest answer is 13 months and you were planning to run it for three weeks, that's worth knowing on day one — not after three weeks of watching numbers that were never going to settle.

NavonaAI's experiment setup shows this estimate up front, before you commit to a test, using your store's own cart volume and conversion rate rather than a generic assumption. It won't stop you from running a long-shot test if that's what you want. It just means you'll know it's a long shot before you start, instead of after.

Frequently asked questions