Guide · Testing
How to split-test when you don’t have the traffic
The usual advice for a low-traffic store is not to test. That is half right: do not run purchase-level tests you cannot read. It is wrong about testing in general, because the arithmetic changes once you move the step you judge.
- 5 min read
- Reviewed 6 October 2026
Why purchase-level tests fail on smaller stores
Statistical power depends on the base rate and the number of events. A purchase is a rare event, so telling a realistic effect apart from noise takes more orders per variant than many stores collect in a month.
The failure is not usually visible as a failure. The test runs, the numbers wobble, someone declares a winner, and the effect never shows up in revenue. That is the low-traffic failure mode: not inconclusive results, but confident wrong ones.
The judging-step trade, stated honestly
Judging at add to cart instead of purchase measures a narrower question: did the change make people more likely to add to cart, not did it make the store more money. Those usually move together and occasionally do not.
That is what the holdback is for. The step-level test decides relatively quickly whether the change did anything; the holdback decides slowly whether it was worth money. Neither alone is sufficient, and together they are honest.
Things that quietly invalidate a test
Peeking at a fixed-horizon test and stopping on a threshold crossing: the most common cause of wins that do not replicate.
Sample ratio mismatch, where the realised split differs from the intended one, usually because of redirects, caching or bot traffic. It means the groups were not comparable and no analysis rescues it.
Running two tests on the same surface at once, which makes both results unattributable.
When not to test at all
A control that does nothing when tapped does not need a test. Neither does a broken discount field or a script from an app you uninstalled last year.
Testing is for changes where reasonable people disagree about the outcome. Fixing something broken is not one of those, and testing it only delays the fix.
How Liftable handles a store that cannot read a test
Liftable runs a power check before any test starts, on your store’s own recent sessions, and refuses a test your traffic cannot read within its 21-day maximum. Where it can, a refusal comes with a readable alternative: the same test measured on reaching checkout or on add to cart, with how long that would take. It never switches the metric on its own; you choose.
Tests are read with a sequential method (mSPRT) set to a 5% false-positive rate, checked for sample ratio mismatch on every read, and a shipped change keeps 10% of visitors on the original so revenue can be measured on Shopify orders against them. This is built, and the statistics are checked in simulation. No test has been read to a result on a merchant’s store.
As a rough guide, about 50,000 sessions a month is where a storefront change can be measured in a few weeks. Below that, only larger effects read in that time, and Liftable tells you which side of the line your store is on before you commit to anything.
Testing at the traffic you have
- 01
Work out the detectable effect before you start
Compute the smallest lift your traffic could tell apart from noise at your baseline rate and planned duration. If it comes back far larger than any change you could plausibly make, the test cannot conclude, and you have saved yourself weeks.
- 02
Move the judging step to where the change acts
Judge a product-page change at add to cart and a cart change at reaching checkout. The event rate rises sharply and the detectable effect shrinks with it.
- 03
Use a sequential method so watching is allowed
A fixed-horizon test is invalidated by checking it early and stopping on a good day. A sequential method stays valid under continuous monitoring, so you can stop as soon as the evidence is sufficient rather than waiting out a pre-computed sample.
- 04
Check for sample ratio mismatch before believing anything
If the realised split differs significantly from the intended one, the groups were not comparable and the result is void. Discard it and fix the cause; never adjust the numbers.
- 05
Ship behind a holdback
Keep a small share of visitors on the original after the winner goes live, and compare afterwards. This is what catches novelty effects and decay, and it is what makes a revenue claim defensible.
Sources
- Liftable: what is built today. Liftable’s own account of what the product does, checked against the code on 2 October 2026 and shown on /how-it-works and /compare. Every figure this page gives about Liftable comes from it. No test has run on a merchant’s store and there are no published merchant results yet.
Questions
How much traffic do I need to A/B test?
It depends on the base rate of what you judge, not on sessions alone. At purchase level many stores cannot read a realistic effect within a month. As a rough guide, about 50,000 sessions a month is where a storefront change can be measured in a few weeks; judging at add to cart brings that within reach of smaller stores for larger effects.
Is it cheating to stop a test early?
With a fixed-horizon test, yes: stopping early on a threshold crossing inflates false positives substantially. With a sequential method it is the intended behaviour, because validity is kept under continuous monitoring.
What is a holdback and do I need one?
A small share of traffic kept on the original after the winner ships, compared afterwards to see whether the improvement lasted. You need one before claiming a revenue figure, because tests win for reasons that do not last.
See whether your own store has the problems this guide describes. The scan is free, takes about a minute, and needs no install.
Scan your storeSolutions
Read next
- Testing · 7 minHow to test a fix without breaking your live storeTest through a route that never edits your theme’s code, read the change first, run one test per surface, keep the rollback to one switch. The safe sequence.Read
- Testing · 7 minCan you automate conversion optimisation end to end?Five of the six steps in conversion optimisation can run without a person. The sixth should not, and the difference matters before you buy anything.Read
- Diagnosis · 7 minWhat a good Shopify conversion rate actually isPublished Shopify conversion rate averages disagree, for defensible reasons. The figures, their sources, why they differ, and how to read your own number.Read
- Diagnosis · 5 minTraffic but no sales: how to find out whyHow to diagnose a Shopify store that gets visitors but few orders: the checks to run, from the cheapest to check to the most expensive, and what each one finds.Read
Terms used
Last reviewed against the product on 6 October 2026.
Find what's holding your Shopify store back
Liftable reads your storefront as a shopper would and shows what looks wrong on your own pages, desktop and mobile: friction, slow pages, missing trust and apps doing nothing, with the evidence for each. Free, in about a minute.
Scan your store.
See where your store is losing revenue and what to fix first. It is free, takes about a minute, and needs no install.