tezvyn:

P-value vs confidence interval in an A/B test

AI-drafted, machine-checkedSource: cxl.combeginner
P-value vs confidence interval in an A/B test
WHAT IT TESTS

Frequentist reasoning beyond binary significance.

ANSWER OUTLINE

A p-value gauges evidence against the null; a 95% CI shows plausible effect sizes and precision.

RED FLAG

Calling the CI a 95% probability the true difference is inside.

WHAT THIS TESTS: This question checks whether you treat statistics as decision-making tools or ritualistic thresholds. Interviewers want to see that you know a p-value is a probability about data under a model, not a probability about truth, and that you use confidence intervals to communicate uncertainty and practical importance. Senior roles require translating these concepts into business language without garbling frequentist definitions.

A GOOD ANSWER COVERS: First, define the p-value as the probability of observing a test statistic at least as extreme as the one measured, assuming the null hypothesis of no difference is true. Emphasize that it does not give the probability the null is true or false, nor does it indicate the size of the effect. Second, define the 95 percent confidence interval as the range of effect sizes that are not rejected by a two-sided test at the 5 percent significance level, obtained by repeating the sampling procedure. Explain that the interval width reflects precision: a narrow interval around a small lift means you can localize the effect, while a wide interval means uncertainty remains large. Third, connect the two by noting that if the 95 percent confidence interval excludes zero, the two-sided p-value will be below 0.05, but the interval adds information about direction and magnitude. Fourth, translate to business impact by stating that a statistically significant p-value with a confidence interval tightly bounded near zero might mean the lift is real but too small to justify engineering costs.

COMMON WRONG ANSWERS: The biggest red flag is Bayesian language applied to frequentist objects, such as saying there is a 95 percent probability the true difference lies inside the interval. Another red flag is claiming a p-value measures the probability the null hypothesis is true or the probability the result occurred by chance. Confusing statistical significance with practical significance is also a common failure mode at senior levels, as is ignoring that confidence intervals assume the model and sampling plan are correct.

LIKELY FOLLOW-UPS: An interviewer might ask how you would choose between a one-sided and two-sided test, how you would handle peeking or multiple comparisons, or what you would do if the confidence interval barely excluded zero but the point estimate implied a huge revenue impact. They may also ask how Bayesian credible intervals differ from confidence intervals, or how sample size affects the width of the interval.

ONE CONCRETE EXAMPLE: Suppose your A/B test shows a conversion lift of 2 percent with a 95 percent confidence interval from 0.1 percent to 3.9 percent and a p-value of 0.04. You should report that the data are incompatible with zero lift at the 5 percent level, but the plausible range spans from a negligible 0.1 percent to a meaningful 3.9 percent. If the feature requires heavy maintenance, you might recommend a larger test before launch rather than celebrating the p-value alone.

Source: CXL

Read the original → cxl.com

Get five bites like this every day.

Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.