What is Simpson's Paradox and how can it bias A/B tests?
Tests whether you recognize that aggregate trends can reverse within subgroups. A strong answer defines the paradox, gives an A/B example where treatment wins overall but loses in every segment due to skewed allocation, and prescribes stratified analysis.
What's really being asked
This question tests whether you understand that aggregate correlation or treatment effects are not causation. Interviewers want to see that you know a statistically significant lift in an A/B test can be completely misleading when subgroups are pooled inappropriately. The core issue is confounding, not random noise. They are looking for statistical maturity and the habit of slicing data before trusting a top-line number.
The full answer
First, a crisp definition: Simpson's Paradox occurs when a trend appears in separate groups of data but disappears or reverses when those groups are combined. Second, a concrete A/B testing scenario, such as a new checkout flow where Variant A wins overall but actually loses to Variant B on every device type because one variant received far more traffic from a high-converting segment. Third, an investigation plan: stratify by suspected confounders like device, traffic source, or time of enrollment; compare segment-level rates against the aggregate; check for unequal sample sizes or selection bias; and apply a regression model or propensity weighting to control for the confounder and recover the true causal effect.
The mistakes people make
A major red flag is attributing the reversal to small sample size or statistical noise without naming a confounding variable. Another is suggesting you simply pick the subgroup results and ignore the aggregate without explaining why the groups differed in size or composition. Saying you would rerun the test longer is also weak unless you explain how you would change the randomization or targeting to balance the confounder. Finally, confusing Simpson's Paradox with the multiple comparisons problem or peeking shows a shallow understanding.
What usually comes next
An interviewer might ask how Simpson's Paradox relates to causal inference and whether stratification alone is enough to fix it. They might probe whether you would ship the feature based on the aggregate win or the subgroup losses, which tests your product judgment and ethical stance on user harm. They could also ask how you would design the experiment upfront to prevent this, such as using stratified randomization or covariate-adaptive randomization to ensure balanced allocation across known confounders.
A concrete example
Imagine an A/B test of a new payment page. Variant A is shown to 1000 mobile users converting at 10 percent and 2000 desktop users converting at 40 percent. Variant B is shown to 2000 mobile users converting at 12 percent and 1000 desktop users converting at 50 percent. Variant A wins overall with 30 percent conversion while Variant B has roughly 24.7 percent, yet Variant B outperforms Variant A on both mobile and desktop individually. The confounder is device type and the allocation is severely skewed: Variant A got mostly desktop traffic with higher baseline conversion, inflating its aggregate rate. To investigate, you would immediately tabulate conversion by device, observe the skewed allocation, and either control for device in a logistic regression or reweight the segments to estimate the true treatment effect.
Interview question
An A/B test shows Variant A significantly outperforms Variant B overall, yet Variant B wins within every user segment. What explains this reversal?
- a.Repeated peeking at the results inflated the false-positive rate for the aggregate metric.
- b.The experiment suffered from insufficient statistical power due to a small sample size.
- c.The variants received unequal proportions of users from segments with different baseline conversion rates.Correct
- d.Segment-level comparisons introduced a multiple-comparisons problem that invalidates the subgroup findings.
Why? this is the answer
This is Simpson's Paradox: skewed allocation across segments with different baseline rates creates a confounded aggregate that reverses the true segment-level effect. Blaming small sample size is wrong because the reversal is structural, not random noise or a power issue.
Just read this? Test yourself on what you have been reading.
Read the original → en.wikipedia.org
- #simpson's paradox
- #a/b testing
- #causal inference
- #confounding
- #experimentation
You just looked this up. Could you explain it out loud?
That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.
The iPhone app is on the way
We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.
Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.
We are hiring for this. Every open role lists the topics its interview covers, so you can prepare for the real thing rather than guessing.
See open roles