More in AI & ML — page 6
Evaluating image generation: FID and IS
WHAT IT TESTS: knowledge of generative image metrics. OUTLINE: FID compares feature distributions of real and generated images, lower is better; Inception Score rewards confident, diverse classes but ignores real data.
Designing an agent that resolves ambiguity
WHAT IT TESTS: agent design for under-specified requests. OUTLINE: detect ambiguity, gather evidence with the contact API, resolve relative time deterministically, ask the user only when genuinely uncertain, then confirm before the irreversible booking.
Securing tool-using LLM agents
WHAT IT TESTS: threat modeling for agentic LLMs. OUTLINE: name indirect prompt injection, data exfiltration, and unsafe tool execution, then defend with sandboxing, least-privilege scoped tools, input/output filtering, and human-in-the-loop on risky actions.
Hybrid search and re-ranking for retrieval
WHAT IT TESTS: knowledge of retrieval beyond plain vectors. OUTLINE: hybrid search fuses dense semantic and sparse keyword signals to catch exact terms dense misses; a cross-encoder re-ranker rescoring top-k boosts precision.
Evaluating a RAG system end to end
WHAT IT TESTS: ability to separate retrieval and generation quality. OUTLINE: measure retrieval with context recall or precision, and generation with faithfulness and answer relevance, attributing failures to the right stage.
Direct Preference Optimization explained
WHAT IT TESTS: understanding of DPO versus RLHF. OUTLINE: DPO reparameterizes the RLHF reward in terms of the policy itself, turning alignment into a simple classification loss on preference pairs with no separate reward model or PPO.
Reward models in RLHF and PPO
WHAT IT TESTS: understanding of the reward model in RLHF. OUTLINE: it learns from human preference comparisons to score responses, then supplies the reward signal that PPO maximizes while a KL penalty keeps the policy near the reference.
Pre-training versus fine-tuning an LLM
WHAT IT TESTS: grasp of the two-stage LLM training lifecycle. OUTLINE: pre-training is broad self-supervised next-token prediction on huge corpora at massive cost; fine-tuning adapts on small labeled data cheaply.
Demographic Parity versus Equalized Odds in hiring
WHAT IT TESTS: understanding fairness definitions. OUTLINE: demographic parity equalizes selection rates regardless of qualification; equalized odds equalizes true and false positive rates across groups, conditioning on the true label.
Explain an interaction effect to a non-statistician
WHAT IT TESTS: communicating interaction effects plainly. OUTLINE: define interaction as it depends on, show separate slope lines per age group, give the business takeaway on targeting.
Present a small but significant A/B test lift
WHAT IT TESTS: structuring an experiment narrative. OUTLINE: hypothesis, design and validity checks, result with effect size and interval, business impact of 0.5%, then a clear recommendation.
Interactive versus static plots for EDA
WHAT IT TESTS: matching viz tooling to the task. OUTLINE: interactive libraries win for exploring dense, high-cardinality, or multi-dimensional data via zoom, hover, and filtering; static plots win for reproducible, publication output.
Parquet versus CSV for analytical data lakes
WHAT IT TESTS: columnar versus row storage trade-offs. OUTLINE: Parquet stores by column enabling projection pushdown, compression, and predicate skipping; CSV is row-based, untyped, and slow to scan.
Audit an ML pipeline for GDPR compliance
WHAT IT TESTS: applying GDPR principles technically. OUTLINE: inventory data and check minimization, verify processing matches stated purpose, build lineage to trace any prediction's inputs.
A/B test two fraud models in production
WHAT IT TESTS: production model experimentation design. OUTLINE: randomize by entity, consider shadow mode first, collect precision/recall and business loss, decide with significance and guardrails.
Communicate a forecast interval to an executive
WHAT IT TESTS: communicating uncertainty to leadership. OUTLINE: give the point estimate but frame the range as scenarios, use a fan chart, tie the interval to planning decisions and risk. RED FLAG: presenting $10M as a guaranteed single number with no range.
Purpose of watermarks in Spark Structured Streaming
WHAT IT TESTS: streaming state management. OUTLINE: a watermark sets a threshold on event-time lateness, lets late data update windows up to that bound, and tells Spark when to finalize and drop old state. RED FLAG: confusing event time with processing time.
repartition() versus coalesce() in Spark
WHAT IT TESTS: Spark partition control. OUTLINE: repartition does a full shuffle and can increase or balance partitions; coalesce avoids a full shuffle and only reduces them. RED FLAG: thinking coalesce can increase partitions or always beats repartition.
Explain a loan denial with LIME or SHAP
WHAT IT TESTS: local explainability and its limits. OUTLINE: LIME fits a local surrogate, SHAP attributes the prediction across features via Shapley values, both give per-feature contributions.
Design an automated A/B test reporting system
WHAT IT TESTS: scalable experiment reporting design. OUTLINE: standardized metric definitions, automated stats with confidence intervals and guardrails, segment breakdowns, a clear ship recommendation.