Skip to content
tezvyn:

A key metric dropped 20%. How would you investigate?

Source: newrelic.comHardHow cards are made

A key metric dropped 20%. How would you investigate?

This tests systematic diagnosis of critical issues. A great answer segments the drop (by region, platform), then traces data upstream from the dashboard to the source, correlating with technical metrics. A red flag is jumping to code before scoping the impact.

What's really being asked

This question assesses your structured problem-solving skills under pressure. The interviewer wants to see if you can create a logical, prioritized investigation plan that moves from broad to specific. It tests your understanding of the entire data lifecycle, from client-side event generation to dashboard aggregation, and your ability to differentiate a technical failure from a real-world business event. It's a test of leadership and diagnostic methodology, not just technical knowledge.

The full answer

A strong answer outlines a phased approach. First, CONTAIN and COMMUNICATE: acknowledge the issue and inform stakeholders. Second, SCOPE and SEGMENT: break down the 20% drop. Is it across all users, or specific to iOS users, a certain region, or a new feature? This is the most critical step. If only Android revenue is down, the problem is localized. Third, TRACE UPSTREAM: start at the dashboard query and work backwards. Check the aggregation layer, the ETL/ELT jobs, the raw data lake/warehouse tables, and finally the source systems (e.g., application databases, event streams like Kafka). Look for changes in data volume, schema, or format. Fourth, CORRELATE with technical metrics: use observability tools. Did error rates spike? Did latency increase for a specific service like a payment processor? Did traces for successful purchase events start showing failures in a new microservice?

The mistakes people make

A major red flag is jumping to a specific technical cause without evidence, for example, immediately saying "I'd check the deployment logs." This ignores the possibility of a client-side bug, a third-party outage (e.g., a payment gateway), a data pipeline schema change, or a real business event. Another weak answer focuses only on the data pipeline, ignoring the application services that generate the data. A senior candidate must think end-to-end. Failing to mention segmentation is the most common mistake; it shows a lack of a systematic approach.

What usually comes next

"Okay, you've segmented the drop to iOS users on the latest app version. What next?" (Tests drilling down). "How would you confirm it's a data bug versus a real drop in user purchasing on iOS?" (Tests validation techniques). "What long-term fixes would you propose to prevent this class of issue?" (Tests proactive thinking, e.g., data contracts, anomaly detection monitors).

A concrete example

To differentiate a bug from a business downturn, you compare two metrics. If "Revenue" is down 20% but "Add to Cart" events are stable, it points to a problem in the checkout or payment flow (likely a technical bug). However, if both "Revenue" and "Add to Cart" are down 20%, it suggests a broader issue with user intent or traffic acquisition (more likely a real business downturn, perhaps due to a competitor's sale or the end of a marketing campaign). This comparative analysis is key.

Interview question

A key business metric suddenly drops 20%. According to structured diagnostic principles, what is the most effective initial step to take?

  • a.Segment the metric by dimensions like platform, region, or app version to scope the impact.Correct
  • b.Validate the data pipeline's integrity by checking ETL job statuses and raw data volumes.
  • c.Alert stakeholders and immediately begin reviewing the dashboard's underlying SQL query.
  • d.Check recent code deployments and server error logs for potential causes.
Why?

The most effective initial step is to segment the drop to understand its scope (e.g., is it only on iOS?). This systematically narrows the problem space before jumping to specific technical causes like code or data pipelines.

Just read this? Test yourself on what you have been reading.

Read the original → newrelic.com

You just looked this up. Could you explain it out loud?

That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.

The iPhone app is on the way

We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.

Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.

Get it on Google PlayiPhone app coming soon

We are hiring for this. Open roles that interview on data — each one lists the topics its interview covers.

See open roles