Confidence intervals versus p-values, explained simply
statistical literacy for A/B decisions.
define a CI as a plausible range for the true effect with its precision, contrast it with a binary p-value, and read [-1%, 7%] as inconclusive because it spans zero.
WHAT THIS TESTS The interviewer checks both correct statistical understanding and the ability to translate it into a clear product decision, since misreading intervals leads to shipping non-effects.
A GOOD ANSWER COVERS Define a confidence interval as a range of plausible values for the true effect, computed so that the procedure captures the true value in ninety-five percent of repeated experiments. Crucially it conveys both the direction and magnitude of an effect and how precise the estimate is: a narrow interval means a well-pinned estimate, a wide one means high uncertainty. Contrast this with a p-value, which reduces everything to a single threshold answer about whether the effect differs from zero, telling you nothing about how big or how precise. Then read the given interval: [-1%, 7%] includes zero, so the result is not statistically conclusive; the true effect could be a small negative, nothing, or up to seven percent positive. Its width signals the sample is too small to decide, so the honest recommendation is to gather more data rather than ship on a hunch.
COMMON WRONG ANSWERS Stating there is a ninety-five percent probability the true value lies in this specific interval, which is a frequentist misinterpretation. Treating any positive point estimate as a win despite the interval crossing zero. Saying a non-significant result proves no effect, when it may just be underpowered.
LIKELY FOLLOW-UPS How does sample size change interval width. What is the difference between statistical and practical significance. How would you decide a minimum detectable effect up front. Why can a p-value below the threshold still represent a trivial effect.
ONE CONCRETE EXAMPLE To the product manager: our best guess is a three percent lift, but the data is consistent with anything from a one percent drop to a seven percent gain, so we cannot confidently say the change helps. Because the range crosses zero, we should not ship yet; running the test longer to collect more samples will narrow the range and let us decide.
Read the original → en.wikipedia.org
Get five bites like this every day.
Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.