How would you determine if Feature X causally drives higher retention?
Tests causal inference intuition for product metrics. Great answers propose a randomized holdback or instrumental variable, control for user intent, and estimate a local average treatment effect.
WHAT THIS TESTS: This question tests whether you understand that product usage is almost never randomly assigned. Users who discover and adopt a feature are usually more motivated, more familiar with the product, or on a different onboarding path than those who do not. The interviewer wants to see if you default to experimental design rather than trusting raw observational metrics, and whether you can articulate why selection bias makes the 20% figure unactionable on its own.
A GOOD ANSWER COVERS: First, acknowledge the self-selection problem: the 20% lift likely mixes the true effect of the feature with the fact that already-happy users are the ones who find it. Second, propose a randomized experiment, such as a feature holdback where eligible users are randomly denied access to measure the counterfactual. Third, if a true experiment is impossible, suggest a natural experiment like an instrumental variable, for example using a UI prompt or eligibility rule that affects feature exposure but not retention directly. Fourth, mention checking pre-period retention or activity levels to control for baseline engagement as a confounder. Fifth, describe the follow-up analysis: an intent-to-treat estimate in the experiment, or a difference-in-differences model if using observational data with a clear before-and-after.
COMMON WRONG ANSWERS: The biggest red flag is saying the correlation is enough to justify a full rollout. Another weak pattern is proposing to control for every variable in a regression without explaining why unobserved confounders might still bias the result. A third mistake is ignoring the practical question of sample size or experiment length, which matters because a null result from an underpowered test is not evidence of no effect.
LIKELY FOLLOW-UPS: The interviewer might ask how you would design the holdback if denying the feature creates user complaints, or how you would interpret the result if the experiment shows no lift but the observational data still shows 20%. They might also ask how you would handle network effects if the feature is social, or whether you would look at heterogeneous treatment effects across user segments.
ONE CONCRETE EXAMPLE: Suppose Feature X is an advanced export button. You notice exporters retain 20% better. Instead of shipping it to everyone, you run a two-week experiment where 10% of eligible users never see the button. If the holdback group retains at the same rate as the exposed group, the feature is not causal; the lift came from power users self-selecting. If the holdback retains 18% lower, you have strong evidence of causality and can estimate the true treatment effect is roughly that gap. You would then follow up by analyzing whether the effect is consistent across weekly versus monthly active users.
Read the original → en.wikipedia.org
Get five bites like this every day.
Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.