What are your null and alternative hypotheses for this A/B test?

This tests translating a directional business question into statistical hypotheses. A strong answer states H0 as no improvement (or a point-null equality for a specified test) and H1 as green outperforming blue. A red flag is framing H0 as 'blue is better' or changing the tail after seeing data.
What's really being asked
The interviewer wants to know if you can move beyond vague intuition and lock a business question into a formal statistical framework. Specifically, they are checking whether you can define a status-quo or no-improvement null, recognize the directional alternative because the prompt asks whether green increases registrations, and name the metric and population rather than speaking generally about button colors.
The full answer
First, define p_green and p_blue as registration conversion rates for the target population. For the directional question, state H0: p_green <= p_blue and H1: p_green > p_blue. A test may use the point null p_green = p_blue to calculate its reference distribution, but the null represents no improvement or a worse outcome as well. Third, explicitly name the exact metric such as registration conversion rate rather than just saying signups, name the two variants being compared, and name the population such as all users landing on the page during the experiment period. Fourth, optionally note that this setup implies a one-tailed test because the company only cares about an increase, not simply any difference.
The mistakes people make
A major red flag is stating the null hypothesis as the blue button is better or that the green button does not work; the null can include no improvement and outcomes worse than the comparison; a point-null equality is a common reference value. Another red flag is switching from a pre-specified directional test to a two-tailed or one-tailed test after seeing the data. Choose the tail in advance and use a two-tailed test when either direction matters. A third red flag is forgetting to define the metric or population, for example saying the null is that the buttons are the same without clarifying what same means in terms of measurable user behavior.
What usually comes next
The interviewer may ask how you would choose between a one-tailed and two-tailed test in practice, or what happens to your error rates if you switch to a two-tailed test after seeing the data. They might ask you to define the practical significance threshold, such as the minimum increase in registration rate that would justify a full rollout, or they might pivot to asking how you would randomize users and what sample size you need to achieve adequate power.
A concrete example
Suppose the current blue button has a registration conversion rate of 12 percent. For a directional test, define H0 as p_green <= p_blue (no improvement over the blue baseline) and H1 as p_green > p_blue for the population visiting the signup page during the two-week experiment. A point-null value of p_green = 12 percent can define the reference distribution for a particular test. Choose this direction before inspecting results, then interpret the p-value alongside the effect size and practical significance.
Interview question
Which hypothesis pair best tests whether a green signup button increases registrations compared to blue?
- a.H0: There is no difference in registration conversion rate; H1: The registration conversion rates are not equal.
- b.H0: The registration conversion rate for green equals that for blue; H1: The registration conversion rate for green is strictly greater than that for blue.Correct
- c.H0: The blue button has a higher registration conversion rate than the green button; H1: The green button has a higher registration conversion rate than the blue button.
- d.H0: The green and blue buttons perform the same; H1: The green button performs better.
Why? this is the answer
The correct answer frames H0 as no difference and H1 as a directional increase in the specific metric, matching the one-tailed nature of the business question. Option A is tempting because it uses 'no difference,' but it wrongly uses a two-tailed alternative that ignores the directional ask and wastes statistical power.
Just read this? Test yourself on what you have been reading.
Read the original → en.wikipedia.org
- #ab-testing
- #hypothesis-testing
- #data-science
- #statistics
- #experimentation
You just looked this up. Could you explain it out loud?
That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.
The iPhone app is on the way
We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.
Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.
We are hiring for this. Open roles that interview on ab-testing — each one lists the topics its interview covers.
See open roles