The statistics, in the open
How Liftable measures
Every number in the product comes from the rules on this page, and the page is rendered from the same constants the engine runs on. If you have a data team, hand them this.
- alpha 5%
- power 80%
- 21-day maximum
- 10% holdback for 30 days
- Reviewed 6 October 2026
What is registered before a test starts
Every test stores a pre-registration before any traffic is split: the metric it will be read on, the minimum detectable effect it is sized for, the false-positive rate (alpha, 5% by default) and the guardrails that can stop it. The default metric is conversion rate at a 20% relative effect; revenue per visitor at 5% when the pre-registration names it. A result is only ever read against the design that was registered.
Add-to-cart rate and checkout-start rate are offered as proxies for a store that cannot read revenue in time. A proxy test carries a revenue guardrail, never ships automatically and is never shown as verified revenue. The page labels it as a proxy.
The power check, and why it refuses
Before a test starts, the engine takes the store’s last 28 days of sessions on that surface and the metric’s variance, and computes how long it would take to detect the registered effect at 80% power. If that is longer than the limit the store set, the test is refused and the page says how many sessions it would need. A sequential design pays for the right to look continuously with a larger sample: the plan allows 2× the fixed-horizon sample, a factor measured in simulation rather than taken from a textbook.
A refusal suggests a readable design at the same relative effect: conversion rate first when the refused test was revenue, and a proxy only when neither reads in time. Never starting a test the store cannot read is the rule; the refusal is the product doing its job.
The test itself
Each test runs a mixture sequential probability ratio test on the difference in the arms’ rates. Its p-value is always valid: a cron job can read it every hour without inflating the false-positive rate, which is what happens when a fixed-horizon test is read daily and stopped on a good morning. The test enters its reading stage when the likelihood ratio crosses the registered boundary, when the maximum duration is reached, or when a person stops it.
Checked in simulation: over 1,000 runs on A/A data (no real difference), the false-positive rate must stay within one point of the 5% it is set to, and over 1,000 runs with a real lift at the planned horizon, at least 80% must cross the boundary. Both are acceptance tests in the repository, not claims in copy.
What stops a test
A hard stop at 21 days by default: a test that reaches it without an answer is closed as not called, never nudged over a line it did not cross. Guardrails are armed from the first read: a drop in revenue per session or in contribution margin past the registered stop rule, a rise in checkout errors, a rise in page load time. A stopped test keeps its record and says why it stopped.
A sample-ratio check runs at every read: the observed split between arms is compared with the split the test was configured to serve, and a mismatch below p = 0.001 pauses the test, because in that state no downstream p-value means anything however good it looks. Revenue per visitor is not read at all until each arm has at least 30 orders, below which the arm mean is a handful of skewed draws.
Who counts
An exposure means the browser was actually served its arm, not merely assigned one. Only storefront sessions count: the checkout pixel’s per-checkout rows are never counted as sessions, or every checkout would count twice. A session from a shopper who has not consented to analytics is never in a test, because the pixel never assigned one. On a store with a live order connection, only orders Shopify confirms count as conversions; a pixel-only purchase is a lost webhook or a forgery.
What makes a dollar figure
Winning the test is not a revenue claim. A shipped change keeps a 10% holdback on the old version for 30 days, and the receipt compares real Shopify orders between the two arms. Only that comparison puts a number next to a dollar sign; it is restated when refunds and cancellations arrive, so a figure on screen shrinks when the orders behind it do. A verification window that closes without an answer contributes nothing to the total and says so.
Built: revenue measured on Shopify orders against a 10% holdback, never projected. No store has a verified line on it yet.
Where this has run
Shipping and discount tests run as Shopify Functions, and a shipping test has run on our own development store. Page tests run through Liftable's theme app extension. No test has run on a merchant's store.
A sequential test (mSPRT) with a power check that refuses underpowered tests. Checked in simulation: over 1,000 runs on A/A data, false positives must stay within a point of the 5% it is set to. No experiment has been read to a result on a merchant's store.
Terms used
Last reviewed against the product on 6 October 2026.
Scan your store.
See where your store is losing revenue and what to fix first. It is free, takes about a minute, and needs no install.