5 Signs Your A/B Test Results Are Misleading You
A/B testing dashboards are good at showing you a winner. They're less good at showing you when that winner is an illusion. Here are five warning signs that a result deserves a second look before you act on it.
1. The result became significant right after you started checking daily
If you find yourself refreshing the dashboard every day waiting for significance, and you stop the test the moment it crosses the line, you've likely fallen into “peeking” — a well-documented way to inflate false positives. The more often you check and the more eager you are to stop early, the higher your real chance of declaring a winner that's actually noise.
Fix: Decide your sample size or test duration in advance, and stick to it, rather than stopping the moment the number looks good.
2. The lift looks big, but the sample size is small
A 25% lift sounds dramatic. A 25% lift based on 40 total conversions sounds a lot less dramatic once you realize how few data points that represents — a handful of conversions shifting between groups can swing a percentage like that easily, with no real signal behind it.
Fix: Always look at raw conversion counts alongside the percentage. If the actual number of conversions is in the dozens rather than the hundreds, treat the result as preliminary.
3. The result only holds up in the aggregate, not in any one segment
This is the core blind spot stratification exists to catch. If you can't find a single user segment where the “winning” variant is clearly ahead — but the overall number still looks good — it's possible the result is being driven by a coincidental mix shift in traffic during the test window, not a real effect of the variant itself.
Fix:Before trusting an aggregate result, check whether it holds up across your major segments. If it doesn't hold up anywhere specific, be skeptical of why it shows up overall.
4. The test ran during an unusual time window
A test that ran entirely over a holiday week, during a major marketing campaign, or during a site outage on one of the variants can produce results that say more about that specific time period than about the variants themselves.
Fix: Note anything unusual happening during your test window, and consider re-running the test during a more typical period if anything irregular occurred.
5. You're testing multiple metrics and one of them happened to come out significant
If you're tracking ten different metrics on a single test, basic probability says that even with no real effect anywhere, you'd expect roughly one of them to look “significant” purely by chance at the standard 5% threshold. Teams sometimes unconsciously gravitate toward whichever metric happened to look good and present that one as “the result.”
Fix:Decide your primary metric before running the test, and treat any other metrics that happen to move as directional, not conclusive, unless you've adjusted your significance threshold to account for checking multiple metrics.
The common thread
Every one of these issues comes from the same root cause: a single significant-looking number can hide a lot of context that changes how much you should trust it. Sample size, timing, consistency across segments, and how many things you tested all affect whether a “winner” is something you should actually act on.
None of this means statistics are useless — it means a good result deserves a quick sanity check before you treat it as fact, the same way you'd double check an important number in a financial report before presenting it to your team.
Stratafy surfaces sample size, segment consistency, and confidence trends automatically on every test, so these red flags are visible before you make a decision, not after. See how it works.