When Should You Stratify Your A/B Test? (And When You Shouldn't)

Stratification is a powerful tool, but it's not a default setting you should flip on for every test. Like most analysis techniques, it adds real value in some situations and adds noise without benefit in others. Here's how to tell which one you're in.

When stratification helps

You have a genuinely diverse user base.If your traffic comes from multiple channels, device types, or customer segments with meaningfully different behavior, there's real variance for stratification to uncover. A B2B SaaS product with both self-serve and enterprise customers is a good candidate — those two groups often behave nothing alike.

The decision has real stakes.If shipping the wrong variant means a meaningful revenue or retention hit, it's worth the extra rigor of checking whether the result holds up across segments before you commit. Pricing page tests, checkout flow changes, and core navigation changes all qualify.

You suspect (or have seen before) a segment-specific reaction.If you've been burned before by a result that didn't generalize, or if you have a hypothesis that, say, mobile users will respond differently than desktop users, stratification directly tests that hypothesis instead of leaving it to guesswork.

You have enough traffic to support it. Splitting a test into segments only works if each segment has enough sample size to produce a meaningful result. A test with 200 total visitors split six ways gives you segments too small to say anything statistically reliable.

When you should skip it

Your test population is small or homogeneous.If you're testing on a narrow audience — say, a single landing page with one traffic source and one device type — there's no real variance to stratify, and breaking the (already small) sample into smaller pieces just makes every result noisier.

The decision is low-stakes.Testing a minor copy change on a page few people see doesn't need the same rigor as a pricing decision. Sometimes “good enough, ship it” is the right call, and adding stratified analysis is more process than the decision warrants.

You don't have a clear hypothesis for why segments would differ. Stratifying by every field you happen to have data on, without any reason to think they'd behave differently, is more likely to produce false positives (apparent differences that are just noise) than real insight. Pick segments you have a reason to check, not every segment available.

Your sample size per segment would be too small to trust.This is worth repeating: an underpowered segment-level result is often worse than no segment-level result at all, because it looks precise while actually being a coin flip. If you can't get a meaningful sample size in a segment, don't pretend you have a conclusive answer for it — flag it as inconclusive instead.

A simple rule of thumb

Ask yourself: “If this test were wrong for one specific group of users, would I want to know before I shipped it everywhere?” If the answer is yes and you have enough traffic to check, stratify. If the answer is “it wouldn't really matter” or you don't have the sample size to check reliably, a standard A/B test is the right tool for the job.

Stratification is meant to add confidence and catch blind spots — not to replace judgment about which tests actually need that extra layer of scrutiny.

Stratafy lets you stratify selectively — choose the segments that matter for each test instead of drowning every result in breakdowns you don't need. See how it works.