Skip to content
tezvyn:

Explain the difference between statistical and practical significance

Source: statsig.comMediumHow cards are made

Explain the difference between statistical and practical significance
Summary

If you know p-values detect real effects but ignore value.

Key points

Define statistical vs practical significance; note large samples make tiny effects significant; give a real example.

Watch out for

Using p < 0.05 alone to justify shipping.

What's really being asked

This question checks whether you treat experimentation as a business decision tool rather than a purely mathematical exercise. Interviewers want to see that you know a low p-value only tells you an effect is unlikely to be noise; it does not tell you if the effect is large enough to justify engineering time, maintenance, or user friction. Senior candidates are expected to connect sample size, effect size, and cost-benefit analysis in their reasoning.

The full answer

First, define statistical significance as the probability that the observed effect is real and not due to chance, typically assessed with a p-value threshold like 0.05. Second, define practical significance as whether the magnitude of that effect, the effect size, matters in the real world given the context. Third, explain the sample size trap: with millions of users, even a 0.1 percent lift can reach statistical significance because the standard error shrinks, but the absolute business impact may be tiny. Fourth, describe a clear decision framework that weighs the lift against implementation cost, opportunity cost, and technical debt.

The mistakes people make

A major red flag is saying that a p-value below 0.05 automatically means you should ship the feature. Another weak response is ignoring effect size entirely and only talking about confidence intervals or p-values. Some candidates also conflate practical significance with sample size, arguing that small samples are the only reason to care about practical significance, which misses the point. Finally, giving a vague example without numbers, such as saying a change improved metrics but was not important, sounds hand-wavy; senior answers use concrete percentages and dollar or hour estimates.

What usually comes next

The interviewer may ask how you would set a minimum detectable effect before launching an experiment, or how you balance statistical power with runtime costs. They might also probe whether you would ship a feature with a large effect size that is not statistically significant, which tests your understanding of Type I and Type II errors in business context. Another common thread is how you communicate a statistically significant but practically insignificant result to a product manager who is emotionally invested in shipping.

A concrete example

Imagine you run an A/B test on a checkout flow with ten million users. The variant shows a 0.05 percent increase in conversion with a p-value of 0.01, so the result is statistically significant. However, the variant requires adding a new microservice that costs fifty thousand dollars a year to maintain and adds two hundred milliseconds of latency. The absolute revenue gain from the lift is projected at twenty thousand dollars a year. Because the engineering cost and latency penalty outweigh the tiny gain, the result is not practically significant and you should not ship it.

Interview question

In an A/B test with ten million users, a 0.05% lift has p = 0.01 but requires a new microservice. Why should you likely decline to ship?

  • a.The projected revenue gain is too small to justify the engineering and maintenance costs.Correct
  • b.The p-value means there is only a 1% probability that the lift is real and not noise.
  • c.The result lacks practical significance only because the large sample size makes small effects detectable.
  • d.The large sample size inflated the standard error and created a false positive.
Why?

Practical significance weighs whether the magnitude of a statistically significant effect justifies its real-world implementation and maintenance costs. Option C is tempting but wrong because practical significance is always about business value, not merely an artifact of sample size.

Just read this? Test yourself on what you have been reading.

Read the original → statsig.com

You just looked this up. Could you explain it out loud?

That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.

The iPhone app is on the way

We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.

Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.

Get it on Google PlayiPhone app coming soon

We are hiring for this. Open roles that interview on growth — each one lists the topics its interview covers.

See open roles