A/B Test Results with Skewed Traffic: What's Next?

This tests your ability to spot confounding variables. A good answer invalidates the results due to sampling bias, proposes segmenting the data by device to find the true effect, and suggests re-running the test with correct randomization.
What's really being asked
This question probes your understanding of experimental validity, specifically your ability to identify confounding variables and sampling bias. The interviewer is looking for you to recognize that the unequal distribution of mobile users makes a direct comparison of the overall metric lift invalid. This scenario is a classic example of Simpson's Paradox, where a trend that appears in aggregate data disappears or reverses when the data is broken down into subgroups. It tests if you prioritize statistical rigor over misleading top-line numbers.
The full answer
A senior-level answer will address four key points in order. First, state clearly that the 5% lift is not trustworthy due to the biased sample; the variable 'device type' is a confounding factor. Second, explain WHY it's a problem: mobile and desktop users likely have different baseline behaviors, and the 'lift' might just be the effect of having more mobile users, not the feature itself. Third, propose an immediate investigative step: segment the results. Analyze the metric lift for mobile users (treatment vs. control) and desktop users (treatment vs. control) separately. This is a form of post-stratification. Fourth, recommend the definitive solution: discard the results and re-run the experiment with proper randomization to ensure the device split is statistically identical across both groups.
The mistakes people make
A major red flag is accepting the 5% lift and suggesting a launch. Another common mistake is to simply say 'the results are biased' without offering a concrete, multi-step plan. A more junior answer might suggest trying to re-weight the results to 'fix' them. While post-stratification analysis is a good first step for investigation, it can't fully rescue a fundamentally flawed experiment setup. The primary recommendation must be to re-run the test correctly. Don't get bogged down in the specific math of re-weighting unless asked.
What usually comes next
'What if you can't re-run the test? How would you make a decision?' (Answer: Use the segmented analysis, but present it with heavy caveats about the low confidence and the risk of making the wrong call). 'How would you ensure this doesn't happen again?' (Answer: Improve the randomization logic in the experimentation platform, add automated checks for Sample Ratio Mismatch (SRM) on key demographic dimensions, and establish a pre-launch checklist for experiments).
A concrete example
Imagine the feature is a new checkout button. Let's say mobile users naturally convert at 2% and desktop users at 10%. The control group (50/50 split) has a blended conversion of (0.5 2%) + (0.5 10%) = 6.0%. The treatment group (80/20 split), even with ZERO feature effect, would have a blended conversion of (0.8 2%) + (0.2 10%) = 3.6%. In this case, the skewed traffic shows a massive DECREASE in the metric, highlighting how the composition effect can create misleading results in either direction. The 5% lift in the original problem could be entirely due to mobile users having a higher baseline for that specific metric, not the feature itself.
Interview question
An A/B test shows a 5% metric lift, but the treatment group has a much higher percentage of mobile users. What is the most appropriate immediate action?
- a.Immediately discard the results and start a new test with proper randomization.
- b.Re-weight the groups using post-stratification to create a corrected overall lift metric.
- c.Segment the results by device type to analyze the lift for mobile and desktop users separately.Correct
- d.Trust the 5% overall lift and recommend launching the feature to capture the gains.
Why? this is the answer
The correct action is to segment the results to understand the feature's true impact on each device group, as the overall lift is likely misleading. Discarding the results without this analysis (Option A) misses a key learning opportunity, even though a re-run is the ultimate solution.
Just read this? Test yourself on what you have been reading.
Read the original → statsig.com
You just looked this up. Could you explain it out loud?
That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.
The iPhone app is on the way
We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.
Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.
We are hiring for this. Open roles that interview on a/b testing — each one lists the topics its interview covers.
See open roles