tezvyn:

Explain Regression Discontinuity Design and propose a real-world scenario

AI-drafted, machine-checkedSource: Wikipedia: Regression discontinuity designadvanced
WHAT IT TESTS

Causal inference via sharp cutoffs when randomization is impossible.

ANSWER OUTLINE

Compare units just above and below a threshold for local effects; propose scenario with forcing variable.

RED FLAG

Calling it randomized or ignoring bandwidth.

WHAT THIS TESTS: This question tests whether you can reason about causal inference without randomization. Interviewers want to see that you understand how to recover a credible treatment effect by exploiting a discontinuous rule that assigns treatment based on a continuous forcing variable. At the senior level, they care about your ability to distinguish local identification from global claims, to discuss practical threats like manipulation or sorting around the cutoff, and to connect the method to a realistic business or policy problem.

A GOOD ANSWER COVERS: First, a precise definition of RDD as a quasi-experimental design that compares units just above and just below a sharp cutoff in a forcing variable to estimate the local average treatment effect at the threshold. Second, a clear explanation of why this works: units near the cutoff are assumed to be otherwise similar, so differences in outcomes are attributed to the treatment. Third, a concrete real-world scenario with a specific forcing variable, cutoff, and outcome, such as a government grant awarded to firms with fewer than fifty employees where the outcome is subsequent hiring. Fourth, a mention of bandwidth selection and the tradeoff between bias and variance, plus a note that you would check for manipulation by testing the density of the forcing variable around the cutoff.

COMMON WRONG ANSWERS: A common red flag is describing RDD as if it were a randomized experiment; the design relies on continuity assumptions, not randomization. Another mistake is proposing a scenario with a fuzzy or self-selected cutoff without explaining how you would handle imperfect compliance or endogenous sorting. Candidates also err by ignoring the local nature of the estimate and claiming the result applies to all units rather than just those near the threshold. Finally, forgetting to discuss balance checks or placebo tests around the cutoff signals shallow familiarity.

LIKELY FOLLOW-UPS: An interviewer might ask how you would test for manipulation of the running variable, such as using a McCrary density test. They could ask about the difference between sharp and fuzzy RDD, or how you would choose a bandwidth and justify a triangular versus rectangular kernel. You might also be asked how you would interpret a null result, or how you would handle a situation where multiple covariates jump at the cutoff, violating the continuity assumption.

ONE CONCRETE EXAMPLE: Suppose a university offers a merit scholarship to every student who scores seven hundred or above on a standardized math exam. You want to estimate the scholarship's effect on first-year GPA. Because students who score six ninety-nine and seven zero one are likely very similar in ability, you compare their outcomes. You would run a local linear regression on either side of the seven hundred cutoff, choose a bandwidth of perhaps ten to twenty points using an Imbens-Kalyanaraman data-driven approach, and test whether the density of test scores is smooth at the threshold to rule out manipulation. The resulting estimate is a credible local average treatment effect for students near the eligibility margin.

Read the original → en.wikipedia.org

Get five bites like this every day.

Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.