Skip to content
tezvyn:

Explain Simpson's Paradox with a user engagement example

Source: statsig.comMediumHow cards are made

Explain Simpson's Paradox with a user engagement example

Tests if you see beyond aggregate data. Define the paradox, give a numerical example where a feature fails overall but wins in segments (e.g., new vs. returning users), and name the confounding variable. A vague definition without numbers is a red flag.

What's really being asked

This question tests your data literacy and skepticism. The interviewer wants to see if you understand that aggregated metrics can be deceptive and that top-line numbers don't tell the whole story. It's a check for your ability to identify confounding variables and the importance of proper segmentation before making a product decision. It separates candidates who just look at dashboards from those who dig into the data.

The full answer

First, provide a clear definition: Simpson's Paradox is a statistical phenomenon where a trend appears in several different groups of data but disappears or reverses when these groups are combined. Second, explain the core mechanism, which is a lurking or confounding variable that is not evenly distributed across the groups. Third, construct a concrete, numerical example related to user engagement. Finally, show the math that proves the paradox, demonstrating both the aggregate failure and the segmented success.

The mistakes people make

A common mistake is only defining the paradox without a concrete, numerical example. This shows theoretical knowledge but not the ability to apply it. Another red flag is providing a confusing or mathematically incorrect example; the numbers must actually work. A critical error is failing to identify the confounding variable. In a user engagement scenario, the user segment itself (e.g., new vs. returning) is the confounder, as different segments have different baseline behaviors and may be unequally represented in the control and test groups.

What usually comes next

Expect questions like: "How would you design an A/B test to prevent this from happening?" (Answer: Ensure proper randomization and maintain consistent traffic allocation throughout the experiment's lifetime). Or, "Based on your example, should we ship the feature?" (Answer: Yes, because it improved the metric for all user segments; the aggregate view was misleading). They might also ask, "What other confounding variables could affect an experiment like this?" (Answer: Device type, geography, acquisition channel, time of day).

A concrete example

Scenario: We A/B test a new "Quick Add" feature to increase the number of items added to a cart. The success metric is the conversion rate of adding an item.

The Paradox: Aggregate Data: The Control group has a 20% conversion rate, and the Test group has a 19% rate. The feature appears to be a failure.

Segmented Data: New Users: Control's conversion is 10% (100 of 1000 users). Test's conversion is 12% (360 of 3000 users). The feature is a success for new users. Returning Users: Control's conversion is 30% (900 of 3000 users). Test's conversion is 31% (310 of 1000 users). The feature is also a success for returning users.

Why it happened: The Test group had a much higher proportion of low-converting new users (3000 vs 1000 in control). This imbalance in the segment mix—the confounding variable—dragged the test group's aggregate conversion rate down, masking its success within each segment. This could occur if traffic splits were changed mid-experiment.

Interview question

An A/B test shows a new feature has a lower overall success rate, but a higher success rate within every individual user segment. What is the most likely conclusion?

  • a.The user segmentation is incorrect and needs to be redefined to resolve the conflicting data.
  • b.The feature is a failure and should not be launched because the overall success rate is lower.
  • c.The feature is successful; the negative overall result is likely caused by an uneven mix of user segments in the test groups.Correct
  • d.The experiment is invalid and must be re-run because the segmented and aggregate results are contradictory.
Why?

This describes Simpson's Paradox. The feature is a success for all users, but a confounding variable (the uneven mix of segments) makes the aggregate result misleading. The most tempting distractor is to trust the aggregate data, which is the primary mistake the paradox highlights.

Just read this? Test yourself on what you have been reading.

Read the original → statsig.com

You just looked this up. Could you explain it out loud?

That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.

The iPhone app is on the way

We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.

Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.

Get it on Google PlayiPhone app coming soon

We are hiring for this. Open roles that interview on analytics — each one lists the topics its interview covers.

See open roles