Explain how CUPED increases statistical power and required data

Tests ANCOVA variance reduction. Answer: CUPED regresses pre-experiment X on Y, shrinking variance by (1-ρ²); needs pre-randomization prognostic baseline; beats difference scores. Red flag: calling it Y-X subtraction or saying it changes the effect.
WHAT THIS TESTS: Whether you understand that CUPED is a regression-based covariate adjustment technique (ANCOVA) rather than a simple difference score, and whether you can articulate why it reduces variance, what data it requires, and how it compares to naive baseline subtraction.
A GOOD ANSWER COVERS: First, the core mechanism. CUPED constructs an adjusted outcome by regressing a pre-experiment baseline covariate X against the post-experiment outcome Y. The adjusted metric has variance approximately equal to (1 minus rho squared) times the original variance, where rho is the correlation between pre- and post-experiment measurements. Second, the data requirements. You need pre-randomization data on the same metric or a strongly prognostic covariate measured before assignment. The baseline must be prognostic for the outcome and should not be influenced by the treatment. Third, the contrast with difference scores. A naive Y minus X approach has variance proportional to 2 sigma squared (1 minus rho) and is only more efficient than the raw mean when correlation exceeds 0.5. CUPED uses an optimal regression weight, so it is never worse than the unadjusted estimator and is usually better than difference scores at any correlation level. Fourth, practical implications. Because sample size scales with the square of standard deviation, even a modest correlation of 0.5 yields roughly a 25 percent variance reduction, which can halve required sample size or duration.
COMMON WRONG ANSWERS: Calling CUPED simple before-after subtraction. Claiming it changes the expected treatment effect estimate rather than just shrinking standard errors and tightening confidence intervals. Asserting that any pre-experiment data works regardless of correlation; if rho is near zero, variance reduction collapses and you gain nothing. Ignoring the pre-randomization requirement and using post-baseline covariates that could be affected by treatment, which breaks the statistical guarantee. Confusing CUPED with post-stratification or blocking.
LIKELY FOLLOW-UPS: How would you handle missing pre-experiment data for new users? What happens if the pre-experiment metric has a different distribution than during the experiment period? Can you use multiple covariates or only the same metric? When would you NOT use CUPED? How do you explain to a product manager why the unadjusted and CUPED confidence intervals might differ in width but should center on a similar point estimate?
ONE CONCRETE EXAMPLE: Suppose you are running an A/B test on purchase conversion with a baseline conversion rate of 2 percent and an expected lift of 0.1 percent. Without CUPED, the variance is driven by user-level heterogeneity. You pull each user's pre-experiment purchase count from the 30 days prior to randomization. If the correlation between pre- and post-period spend is 0.8, CUPED reduces variance by roughly 64 percent. That means the standard error drops by about 40 percent, and because sample size is proportional to variance, you need only about 36 percent of the original sample to detect the same 0.1 percent lift at the same power and alpha. If you had instead used a simple difference score, the variance would be 2 sigma squared (1 minus 0.8), which is actually worse than the unadjusted variance, illustrating why CUPED dominates naive subtraction.
Source: zetyra.com
Read the original → zetyra.com
- #cuped
- #ab-testing
- #variance-reduction
- #ancova
- #experimentation
Get five bites like this every day.
Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.