Skip to content
tezvyn:

P-value vs confidence interval in an A/B test

Source: cxl.comEasyHow cards are made

P-value vs confidence interval in an A/B test
Summary

Frequentist reasoning beyond binary significance.

Key points

A p-value gauges evidence against the null; a 95% CI shows plausible effect sizes and precision.

Watch out for

Calling the CI a 95% probability the true difference is inside.

What's really being asked

This question checks whether you treat statistics as decision-making tools or ritualistic thresholds. Interviewers want to see that you know a p-value is a probability about data under a model, not a probability about truth, and that you use confidence intervals to communicate uncertainty and practical importance. Senior roles require translating these concepts into business language without garbling frequentist definitions.

The full answer

First, define the p-value as the probability of observing a test statistic at least as extreme as the one measured, assuming the null hypothesis of no difference is true. Emphasize that it does not give the probability the null is true or false, nor does it indicate the size of the effect. Second, define the 95 percent confidence interval as the range of effect sizes that are not rejected by a two-sided test at the 5 percent significance level, obtained by repeating the sampling procedure. Explain that the interval width reflects precision: a narrow interval around a small lift means you can localize the effect, while a wide interval means uncertainty remains large. Third, connect the two by noting that if the 95 percent confidence interval excludes zero, the two-sided p-value will be below 0.05, but the interval adds information about direction and magnitude. Fourth, translate to business impact by stating that a statistically significant p-value with a confidence interval tightly bounded near zero might mean the lift is real but too small to justify engineering costs.

The mistakes people make

The biggest red flag is Bayesian language applied to frequentist objects, such as saying there is a 95 percent probability the true difference lies inside the interval. Another red flag is claiming a p-value measures the probability the null hypothesis is true or the probability the result occurred by chance. Confusing statistical significance with practical significance is also a common failure mode at senior levels, as is ignoring that confidence intervals assume the model and sampling plan are correct.

What usually comes next

An interviewer might ask how you would choose between a one-sided and two-sided test, how you would handle peeking or multiple comparisons, or what you would do if the confidence interval barely excluded zero but the point estimate implied a huge revenue impact. They may also ask how Bayesian credible intervals differ from confidence intervals, or how sample size affects the width of the interval.

A concrete example

Suppose your A/B test shows a conversion lift of 2 percent with a 95 percent confidence interval from 0.1 percent to 3.9 percent and a p-value of 0.04. You should report that the data are incompatible with zero lift at the 5 percent level, but the plausible range spans from a negligible 0.1 percent to a meaningful 3.9 percent. If the feature requires heavy maintenance, you might recommend a larger test before launch rather than celebrating the p-value alone.

Interview question

An A/B test shows a 2.5% conversion lift with a 95% confidence interval from 0.5% to 4.5% and a p-value of 0.03. Which interpretation is correct?

  • a.Because the result is statistically significant, the feature should be launched immediately.
  • b.There is a 95% probability the true lift falls inside the 0.5% to 4.5% interval.
  • c.The p-value means there is only a 3% chance the observed difference happened by random chance.
  • d.The data are incompatible with zero lift at the 5% level, and the interval shows plausible effect sizes between 0.5% and 4.5%.Correct
Why?

A 95% confidence interval identifies effect sizes that are not rejected by the data under the null, so excluding zero aligns with p < 0.05 while also showing magnitude and precision. Option B is wrong because it treats the frequentist interval as a Bayesian probability statement about the true parameter.

Just read this? Test yourself on what you have been reading.

Read the original → cxl.com

You just looked this up. Could you explain it out loud?

That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.

The iPhone app is on the way

We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.

Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.

Get it on Google PlayiPhone app coming soon

We are hiring for this. Open roles that interview on statistics — each one lists the topics its interview covers.

See open roles