Skip to content
tezvyn:

What does a p-value of 0.03 mean in an A/B test?

Source: analytics-toolkit.comMediumHow cards are made

What does a p-value of 0.03 mean in an A/B test?

This tests your practical grasp of statistical significance. A good answer defines p-value (probability of the result if the null hypothesis is true), explains that p=0.03 is significant vs. alpha=0.05, and concludes you can reject the null.

What's really being asked

This question tests your ability to translate a core statistical concept into a practical business decision. Interviewers want to see if you understand not just the academic definition, but the real-world implications and risks. It separates candidates who have memorized a definition from those who have actually run tests and had to justify shipping a feature based on the results. Your answer reveals your grasp of risk, specifically the risk of a false positive (Type I error).

The full answer

A complete answer should cover four key points in order. First, define the p-value correctly: it is the probability of observing your result, or a more extreme one, ASSUMING the null hypothesis is true. The null hypothesis is the default state that your change has no effect. Second, state the comparison: this p-value is compared against a pre-determined significance threshold (alpha), which is typically 0.05. Third, apply it to the specific number: since 0.03 is less than 0.05, the result is statistically significant. Fourth, state the practical conclusion: you can reject the null hypothesis. This means you have sufficient evidence to conclude the new feature has a real effect on the conversion rate, and it wasn't just random chance. You are making a decision with a 3% chance of being a false positive.

The mistakes people make

The most common wrong answer is stating that the p-value is the probability that the null hypothesis is true, or that there's a 97% chance the new feature is better. This is incorrect and a major red flag, as it confuses frequentist and Bayesian concepts. Another error is being too vague, like saying "there's a 3% chance we're wrong." Wrong about what? A good answer is precise about it being the probability of a false positive. Finally, many candidates forget to include the crucial phrase "assuming the null hypothesis is true" in their definition, which invalidates the entire explanation.

What usually comes next

Expect follow-ups like: "What would you do if the p-value was 0.06?" (Fail to reject the null, don't ship). "What's a Type II error?" (A false negative, failing to detect a real effect). "If we run 100 experiments on useless features, how many 'significant' results would you expect with an alpha of 0.05?" (Around 5, testing your understanding of the false positive rate).

A concrete example

We are testing a new checkout button design (variant) against the old one (control). The null hypothesis is that the new design has no effect on the purchase conversion rate. We set our significance threshold (alpha) to 0.05 before starting. After the test, our new design shows a higher conversion rate with a p-value of 0.03. Since 0.03 is less than 0.05, we reject the null hypothesis. We conclude the new button design caused the increase in conversions. The p-value means there was only a 3% probability of seeing this much of an improvement just by random luck if the button actually had no effect.

Interview question

In an A/B test with an alpha of 0.05, if a new feature yields a p-value of 0.03, what is the most accurate interpretation?

  • a.There is a 97% probability that the new feature is genuinely better than the old one.
  • b.It means there is a 3% chance that the new feature is not truly better, despite the observed improvement.
  • c.There is a 3% probability that the null hypothesis (that the new feature has no effect) is true.
  • d.It is the probability of observing a result as extreme as, or more extreme than, the one obtained, assuming the new feature has no actual effect.Correct
Why?

Option D correctly defines the p-value as the probability of observing the data (or more extreme) under the assumption that the null hypothesis is true. Option C incorrectly interprets the p-value as the probability of the null hypothesis itself being true, which is a common misconception.

Just read this? Test yourself on what you have been reading.

Read the original → analytics-toolkit.com

You just looked this up. Could you explain it out loud?

That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.

The iPhone app is on the way

We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.

Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.

Get it on Google PlayiPhone app coming soon

We are hiring for this. Open roles that interview on a/b testing — each one lists the topics its interview covers.

See open roles