tezvyn:

MLOps & Infrastructure

Model deployment, training infra, experiment tracking

258 bites

More in MLOps & Infrastructure — page 3

Explain dynamic batching in inference servers and its trade-off
MLOps & Infrastructure2 min read

Explain dynamic batching in inference servers and its trade-off

WHAT IT TESTS: Inference scheduling and the latency-vs-throughput trade-off. ANSWER OUTLINE: Dynamic batching launches when a time window or max size is met, improving throughput over static batching, but short ones wait for the slowest.

Design a multi-tenant GPU serving system for hundreds of fine-tuned models
MLOps & Infrastructure2 min read

Design a multi-tenant GPU serving system for hundreds of fine-tuned models

Tests GPU memory tradeoffs versus cold-start latency in multi-tenant serving. Strong answers propose tiered CPU staging, predictive pre-warming, and disaggregated prefill and decode. Red flag: keeping all models GPU-resident or ignoring transfer overhead.

Compare Canary and Blue/Green ML deployments and model-specific metrics
MLOps & Infrastructure2 min read

Compare Canary and Blue/Green ML deployments and model-specific metrics

WHAT IT TESTS: Model quality vs infra health in rollouts. ANSWER OUTLINE: Contrast Canary gradual shift vs Blue/Green instant swap; highlight silent failures, data drift, prediction distribution; cite accuracy and calibration.

Expose a trained model as a simple web service
MLOps & Infrastructure2 min read

Expose a trained model as a simple web service

Practical MLOps knowledge from model serialization to serving. Package the model into a standard format, containerize it, expose a REST endpoint behind a load balancer, and add monitoring. A bare Flask server without containers or health checks is a red flag.

How would you version control a 50GB dataset in a CI/CD pipeline?
MLOps & Infrastructure2 min read

How would you version control a 50GB dataset in a CI/CD pipeline?

WHAT IT TESTS: Code and data versioning without breaking CI/CD speed. ANSWER OUTLINE: Contrast Git LFS (simple, but 50GB chokes CI clones) with DVC (git metadata plus S3; enables selective pulls and CI cache). RED FLAG: Storing 50GB binaries in Git.

How would GDPR requirements influence experiment tracking and model management design?
MLOps & Infrastructure2 min read

How would GDPR requirements influence experiment tracking and model management design?

WHAT IT TESTS: designing for compliance as a systems constraint, not an afterthought. ANSWER OUTLINE: immutable data lineage, user exclusion lists, audit logs, versioned explainability. RED FLAG: manual deletion without model unlearning or provenance.

How do you ensure ML experiment reproducibility beyond random seeds?
MLOps & Infrastructure2 min read

How do you ensure ML experiment reproducibility beyond random seeds?

Tests system-level reproducibility through data versioning, environment capture, and pipeline automation. Strong answers cover versioned datasets, containerized dependencies, and immutable experiment logs.

Describe an ML workflow with massive egress fees and re-architecture to mitigate
MLOps & Infrastructure2 min read

Describe an ML workflow with massive egress fees and re-architecture to mitigate

Tests whether you recognize egress spikes when storage and compute cross cloud or region boundaries. Great answers sketch a multi-cloud training pipeline, cite per-GB rates, and propose caching or compute placement. Red flag: suggesting compression alone.

Design a showback or chargeback system for ML infrastructure costs
MLOps & Infrastructure2 min read

Design a showback or chargeback system for ML infrastructure costs

WHAT IT TESTS: Bridging ML telemetry with FinOps for shared GPU storage. ANSWER OUTLINE: Tag workloads to cost centers; define shared-resource formulas; automate reconciliation; use showback. RED FLAG: Using raw cloud bills as attribution without GL mapping.

MLOps & Infrastructure2 min read

Design a near real-time cost visibility system for ML teams

Tests cost attribution across shared ML infrastructure and streaming pipeline design. Strong answers combine billing exports with resource labels, sub-hour aggregation, and anomaly detection for training spikes.

MLOps & Infrastructure2 min read

How do you adapt ML training for spot instance interruptions?

Tests resilience under preemption. Strong answers cover frequent checkpoints to durable storage, SIGTERM handling, idempotent retries with budgets, and compute-state separation. Red flag: saving checkpoints only on local ephemeral disks or solely at epoch end.

Describe a basic lifecycle policy to manage cloud storage costs
MLOps & Infrastructure2 min read

Describe a basic lifecycle policy to manage cloud storage costs

This tests cost optimization via tiered storage and automated expiration. Strong answers list transitions from Standard to IA to Glacier, then deletion after set days, plus retrieval costs. A red flag is using manual scripts instead of native lifecycle rules.

Differences between on-demand, reserved, and spot EC2 instances?
MLOps & Infrastructure2 min read

Differences between on-demand, reserved, and spot EC2 instances?

Tests cost-reliability-commitment tradeoffs for ML infrastructure. Good answers map on-demand to experiments, reserved for production training, and spot to fault-tolerant batch jobs. Red flag: spot for real-time serving or skipping reserved capacity analysis.

MLOps & Infrastructure2 min read

How do you attribute cloud costs to ML projects and implement tagging?

Tests knowledge of resource tagging for cost attribution. A strong answer names provider-specific tags or labels, embeds them in infrastructure-as-code, and activates cost allocation reports.

MLOps & Infrastructure2 min read

Design a defense-in-depth strategy against adversarial evasion on a deployed image classifier

WHAT IT TESTS: Your ability to layer training-time and inference-time defenses for adversarial robustness. ANSWER OUTLINE: Proactive: adversarial training, preprocessing, ensembles.

Design a cryptographically verifiable ML audit trail from dataset to deployment
MLOps & Infrastructure2 min read

Design a cryptographically verifiable ML audit trail from dataset to deployment

Tests cryptographic provenance and tamper-evident ML pipelines. Strong answers cover content-addressed datasets, signed training logs linking code and hyperparameters to model hashes, and deployment signature checks.

What is the wrong and right way to manage ML database secrets?
MLOps & Infrastructure2 min read

What is the wrong and right way to manage ML database secrets?

This tests secret management hygiene for ML pipelines. A strong answer rejects hardcoded secrets and env vars, then proposes AWS Secrets Manager with IAM retrieval, TLS, caching, and rotation. A red flag is suggesting .env files, ConfigMaps, or CLI arguments.

MLOps & Infrastructure2 min read

How would you programmatically monitor a deployed model for demographic bias?

Tests operationalizing fairness beyond static audits. Track group metrics like parity and equalized odds; slice by protected attributes; alert on drift; route violations to review. Red flag: treating fairness as a one-time check versus continuous monitoring.

MLOps & Infrastructure2 min read

Describe securing an automated ML pipeline and CI/CD integration points

Tests ML supply-chain depth versus bolt-on appsec. Strong answers stage checks across CI/CD: dependency scans at build, container and model scans before registry, plus runtime input guards.

Why was this customer denied: global or local explanation?
MLOps & Infrastructure2 min read

Why was this customer denied: global or local explanation?

This tests matching questions to explanation scope. Global methods show overall behavior; local methods explain one prediction. Specific denials need local methods like SHAP. A red flag is using global summaries like permutation importance or PDPs for a case.