When user-level A/B tests get contaminated
recognizing interference that breaks the independence assumption.
network or marketplace spillover violates SUTVA, so randomize by cluster (geo, group, time) and analyze at that level.
WHAT THIS TESTS This checks deep understanding of the stable unit treatment value assumption: that one unit's outcome is unaffected by another's treatment. The interviewer wants you to spot interference, explain why it biases user-level tests, and pick a randomization unit that restores independence, with honest acknowledgment of the cost.
A GOOD ANSWER COVERS First, a concrete contamination scenario. In a social network, a new sharing feature given to treatment users generates content that control users also see, so control outcomes rise too and the measured difference shrinks toward zero. In a two-sided marketplace, boosting demand for treatment buyers consumes shared inventory, hurting control buyers and inflating the apparent treatment effect. Both violate independence: treatment spills across the assignment boundary. The fix is to randomize at a unit that contains the interactions. Cluster randomization assigns whole groups, such as cities, teams, or social communities, to one arm so interacting users share a treatment. Geo testing splits markets; switchback or time-based designs alternate the whole system between treatment and control over intervals, useful when spillover is global. Crucially, the analysis must respect the clustering: variance is computed across clusters, not individuals, because users within a cluster are correlated. That reduces effective sample size and statistical power, the price of removing bias.
COMMON WRONG ANSWERS Claiming user-level tests are always valid. Randomizing by cluster but then analyzing as if every user were independent, which understates variance and produces false significance. Ignoring that fewer clusters means much lower power. Confusing this with simple sample-ratio mismatch.
LIKELY FOLLOW-UPS How do you size a clustered experiment? Power depends on the number of clusters and intra-cluster correlation. When is a switchback better than geo? When effects are short-lived and the system is shared globally. How do you detect spillover after the fact?
ONE CONCRETE EXAMPLE A food-delivery app tests a courier-incentive change. At the user level, treated couriers grab the best orders, leaving control couriers worse off, so the user-level lift is overstated. The team instead randomizes by city: some cities run the incentive, others do not, and they compare city-level metrics. Interacting couriers within a city share one arm, removing the spillover, and the team analyzes across cities, accepting lower power for an unbiased estimate.
Read the original → docs.geteppo.com
Get five bites like this every day.
Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.