Skip to content
tezvyn:

When is an A/B test not feasible, and what is DiD?

Source: statsig.comMediumHow cards are made

When is an A/B test not feasible, and what is DiD?

This tests your grasp of causal inference when randomization isn't possible. Explain a scenario like a state-level launch, introduce Difference-in-Differences (DiD), and state its core parallel trends assumption.

What's really being asked

This question assesses your ability to move beyond standard A/B testing and apply more sophisticated causal inference techniques. It probes whether you can identify situations where randomization is infeasible (e.g., geographic rollouts, marketplace effects, legal constraints) and still devise a method to measure impact rigorously. It separates candidates who only know textbook experimentation from those who can handle real-world data complexities.

The full answer

First, provide a clear scenario where user-level randomization is impossible. A classic example is a feature launched in one geographic area but not another, like rolling out a new algorithm in California but not Texas because of operational constraints. Second, introduce Difference-in-Differences (DiD) as the quasi-experimental solution. Explain that you calculate the change in your metric for the treatment group (before vs. after) and subtract the change in the metric for the control group over the same period. This double-differencing isolates the treatment effect from underlying trends. Third, and most importantly, state the core assumption: parallel trends. This means that, in the absence of the treatment, the two groups would have followed the same trend over time.

The mistakes people make

A major red flag is simply comparing the post-launch metrics of the two groups (e.g., "California's engagement was 5% higher than Texas's"). This is a correlation-causation fallacy that ignores pre-existing differences. Another common mistake is misstating the core assumption, for instance, claiming the groups must have identical metric values before the change, when they only need to have parallel trends. Proposing a simple pre-post analysis on the treatment group alone is also a significant error, as it fails to control for seasonality or other external factors that could affect the metric.

What usually comes next

Expect follow-ups like: "How would you test the parallel trends assumption?" (A good answer is to plot the metric for both groups for a long pre-period and visually inspect if the lines move in parallel). Another is, "What are other potential issues with DiD?" (Mention other simultaneous events affecting only one group, or changes in the composition of the groups over time).

A concrete example

Imagine we launch a new driver incentive program in New Jersey but not in neighboring Pennsylvania. We track driver acceptance rates for 3 months before and 3 months after launch. Pre-launch, NJ rates were 80% and PA were 78%. Post-launch, NJ is 88% and PA is 81%. The simple post-launch difference is 7% (88-81). The DiD calculation is more accurate: the change in NJ is +8% (88-80), while the change in PA (the control for market trends) is +3% (81-78). The estimated causal effect of the program is the difference between these changes: 8% - 3% = a 5 percentage point increase.

Interview question

A team uses a Difference-in-Differences analysis for a feature launched in France, using Germany as a control. What is the most critical assumption for this analysis to be valid?

  • a.The metric in the control group (Germany) must remain stable and unchanged throughout the entire study period.
  • b.Without the new feature, the metric's trend in France would have been parallel to its trend in Germany.Correct
  • c.The metric of interest must have been at the same level in both France and Germany before the launch.
  • d.Any external events, like holidays, must affect France but not Germany during the study period.
Why?

The core assumption of DiD is 'parallel trends', meaning the treatment and control groups would have followed the same trend without the intervention. It is incorrect that the metric's starting levels must be identical, as DiD is specifically designed to account for such baseline differences.

Just read this? Test yourself on what you have been reading.

Read the original → statsig.com

You just looked this up. Could you explain it out loud?

That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.

The iPhone app is on the way

We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.

Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.

Get it on Google PlayiPhone app coming soon

We are hiring for this. Open roles that interview on analytics — each one lists the topics its interview covers.

See open roles