A/B Testing 101 for Product Managers: How to Run a Test You Can Actually Trust
Most A/B tests that "win" don’t survive a second look. A practical guide to freezing your metric before launch, sizing the test correctly, and avoiding the peeking and segmentation traps that turn noise into a false win.

A/B Testing 101 for Product Managers: How to Run a Test You Can Actually Trust
Most A/B tests that "win" don't survive a second look — the sample was too small, the team peeked early and stopped the moment the line turned green, or someone sliced the results into six segments until one came back significant. The mechanics of running a test are easy to learn from a blog post; the discipline that keeps a result from being a false positive is what actually separates PMs who ship real lift from PMs who ship noise. Here's how to run one you can defend.
1. Start with the decision, not the metric
Before you write a single line of the experiment brief, answer one question: what will you actually do differently depending on which variant wins? If the honest answer is "ship the winner either way because we already committed to the redesign," you're not running an experiment — you're running a formality that happens to produce a chart. A real test exists because the decision is genuinely open, and the losing variant would actually get killed. If nothing changes based on the result, don't spend the traffic.
2. Pick one primary metric and freeze it before launch
Every test should have exactly one pre-registered primary metric that determines win or lose, decided before the test goes live — not selected afterward from whichever secondary metric moved. Track a handful of secondary and guardrail metrics for context, but treat them as directional, not decision-making. The moment you let yourself pick the "best" metric after seeing the data, you've turned a controlled experiment into a search for a story, and with enough metrics in play, something will move by chance alone.
3. Size the test before you run it, not after
Run a sample-size calculation before launch: your baseline conversion rate, the minimum lift worth caring about, and your traffic volume determine how many days the test needs to run — and that number is often bigger than PMs expect for anything short of a dramatic change. A test that's underpowered doesn't fail loudly; it just produces a wide, noisy result that you're tempted to squint at and call a win. If the math says four weeks and the roadmap only has one, that's a signal to test a bigger swing, not to launch anyway and hope.
4. Don't stop the moment it looks significant
Checking results daily and stopping the instant p < 0.05 appears is the single most common way teams fool themselves — called "peeking," and it inflates your false-positive rate far above the 5% you think you're protecting against, because you're really running dozens of small tests (one per day you check) and taking the first lucky one. Commit to the pre-calculated sample size and duration before you look, or use a sequential-testing method built to allow early stopping safely. A result that looked strong on day 3 and evaporated by day 14 was never real in the first place.
5. Run full business cycles, and watch for novelty
A test launched on a Tuesday and stopped the following Monday has skipped an entire weekend's worth of different user behavior. Run tests across full weekly cycles at minimum, and be honest that a new UI element often wins in week one purely because it's new — novelty effects fade, so a lift that's real in week one but gone by week three is telling you the change doesn't hold up, not that you should have stopped early to bank the fake win.
6. Segment for understanding, not for a second chance at a win
If the overall result is flat, it's tempting to slice by device, by geography, by new-versus-returning users until one slice turns green — but with enough cuts, chance guarantees you'll find one. Pre-register the one or two segments you have a real hypothesis for (e.g., "we expect this to matter more on mobile because of the layout change"), and treat any segment you didn't predict in advance as a hypothesis for the *next* test, not evidence for this one. A flat overall result with a compelling story in a post-hoc segment is usually noise wearing a narrative.
Final thoughts
The gap between a test that holds up and one that quietly gets walked back a quarter later almost never comes down to tooling — it comes down to discipline: a metric frozen before launch, a sample size calculated instead of guessed, and the willingness to sit on your hands until the pre-committed window closes. That discipline is also what makes a test worth writing up: a lift you can defend under questioning is exactly the kind of result that belongs in a case study, next to the metric you originally set out to move.
Start your portfolio here — free, no credit card.
FAQ
What is the most common mistake product managers make when running A/B tests?
Peeking — checking results daily and stopping the moment the test looks statistically significant. This inflates the real false-positive rate well above 5%, because checking repeatedly is effectively running many small tests and keeping the first lucky one. The fix is to commit to a pre-calculated sample size and duration before looking, or use a sequential-testing method designed for safe early stopping.
How do I decide how long an A/B test should run?
Calculate the required sample size before launch using your baseline conversion rate, current traffic, and the smallest lift that would actually be worth shipping. Divide that by your daily traffic to get the run time, and commit to it — an underpowered test doesn’t fail loudly, it just produces a noisy result that’s tempting to misread as a win.
Should I pick my A/B test’s primary metric before or after seeing the results?
Always before. Pre-register one primary metric that determines win or lose, and treat every other metric you track as directional context only. Selecting the "best" metric after the data comes in turns a controlled experiment into a search for a favorable story, and with enough metrics available, something will move by chance alone.
Is it okay to segment A/B test results by device or user type?
Only if you predicted that segment mattered before the test launched. Slicing results into many post-hoc segments until one turns significant will almost always surface a green result somewhere by chance. Pre-register the one or two segments you have a real hypothesis for, and treat any others as an idea for your next test rather than proof from this one.
Related reading