Ryz Labs/Interview questions/machine learning engineer
Interview questions

Machine learning engineer interview questions (2026)

Questions on training, evaluation, feature stores, serving and drift that separate engineers who ship models to production from those who stop at a notebook.

These 27 machine learning engineer interview questions focus on getting models trained, evaluated, served and kept healthy in production: PyTorch training loops and their failure modes, distributed training, offline versus online evaluation, feature stores, deployment, latency budgets and drift. They are written for senior ML engineers rather than researchers or analysts. At Ryz, ML engineers are evaluated along the same lines: recruiters source people who have owned models in production, candidates complete structured NTRVSTA AI interviews on training, serving and system design, and recruiters review every candidate before and after. AI scores are advisory, and people make the hiring decisions.

How to use these questions

Decide which half of the role matters more. Some ML engineers spend their time on modeling and training; others build the platform and serving path around models. Weight the intermediate section for the first and the MLOps section for the second, and use the senior section for anyone who will own a system end to end.

Ask for specifics from production: the last model they rolled back and why, how they knew it was degrading, what the latency budget was. Engineers who have shipped answer with numbers and incidents. Finish with a practical exercise, since a working training and evaluation script reveals habits that conversation hides.

Fundamentals

Write a minimal but correct PyTorch training and validation loop. What do people commonly get wrong?

model.train()
for x, y in train_loader:
x, y = x.to(device), y.to(device)
optimizer.zero_grad(set_to_none=True)
loss = loss_fn(model(x), y)
loss.backward()
optimizer.step()

model.eval()
with torch.no_grad():
val_loss = sum(loss_fn(model(x.to(device)), y.to(device)).item()
for x, y in val_loader) / len(val_loader)

Common bugs: forgetting model.eval(), so dropout and batch norm behave as in training; computing validation without no_grad, which wastes memory; and accumulating loss tensors instead of .item(), which keeps graphs alive.

What a strong answer shows: They write it fluently and explain train versus eval mode without being asked.

Training loss keeps falling, but validation loss starts rising after epoch 8. What do you do?

The model is overfitting. Keep the checkpoint with the best validation metric (early stopping), then add regularization: weight decay, dropout, data augmentation or a smaller model. Also check whether the validation set matches production data, because a distribution gap looks similar.

What a strong answer shows: They checkpoint on validation and consider data problems, not only regularization knobs.

Why might gradient-boosted trees beat a neural network on tabular data?

Trees handle heterogeneous feature scales, missing values and sharp thresholds well with little tuning, and train fast on CPUs. Neural networks pay off on tabular data mainly when you have very large datasets, high-cardinality categorical features that benefit from embeddings, or need to combine tabular inputs with text or images.

What a strong answer shows: They start with a strong baseline like LightGBM or XGBoost before reaching for deep learning.

What does the learning rate control, and how do you choose a schedule?

It sets the step size of each update. Too high and loss diverges or oscillates; too low and training crawls or stalls. A common recipe is AdamW with linear warmup followed by cosine decay. A short learning-rate range test helps find a sensible peak.

What a strong answer shows: They explain why warmup stabilizes early training, especially for transformers.

What is an embedding, and how would you use one outside of NLP?

An embedding is a dense vector learned so that similar items are close together. In recommendations, user and item embeddings from a two-tower model enable fast candidate retrieval with approximate nearest neighbor search. In fraud, merchant embeddings capture behavior that one-hot encoding cannot.

What a strong answer shows: They connect embeddings to a concrete retrieval or feature use case.

Batch normalization versus layer normalization: why do transformers use layer norm?

Batch norm normalizes each feature across the batch, so it depends on batch statistics and behaves differently at inference. Layer norm normalizes across features within one example, which works with variable sequence lengths and small or per-token batches. Most modern transformers also place the norm before attention and MLP blocks for stability.

What a strong answer shows: They tie the choice to batch-size dependence and inference behavior.

How do you split data for a model that will predict on future events?

Split by time: train on the past, validate on the following period and test on the most recent. Group by entity if the same user can appear in both sets. Touch the test set once, at the end, or it stops being a test.

What a strong answer shows: They mirror production conditions in the split.

Intermediate

Loss turns to NaN at step 12,000 in a mixed-precision run. How do you debug it?

Check for a learning rate spike, exploding gradients, a bad batch (divide by zero, log of zero) and fp16 overflow. Log gradient norms, add clipping, prefer bf16 on hardware that supports it, and if using fp16, make sure a gradient scaler is in place.

scaler = torch.amp.GradScaler("cuda")
with torch.autocast(device_type="cuda", dtype=torch.float16):
loss = loss_fn(model(x), y)
scaler.scale(loss).backward()
scaler.unscale_(optimizer)
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
scaler.step(optimizer)
scaler.update()

What a strong answer shows: They know to unscale before clipping and can reproduce the failing batch from a checkpoint.

GPU utilization sits at 30% during training. Where is the bottleneck?

Usually the input pipeline. Profile with the PyTorch profiler. Fixes include more DataLoader workers, pin_memory=True, preprocessing offline, faster formats, larger batches and moving augmentation to the GPU. torch.compile can reduce Python overhead once the GPU is actually busy.

What a strong answer shows: They measure before tuning and know the data loader is the usual suspect.

When do you use DistributedDataParallel versus FSDP?

DDP replicates the full model on each GPU and averages gradients; it is simple and fast when the model, gradients and optimizer state fit on one device. FSDP shards parameters, gradients and optimizer state across GPUs, which is needed for models that do not fit, at the cost of more communication.

What a strong answer shows: They reason about memory per GPU, including optimizer state.

A new model wins offline but the online A/B test is flat. Why might that happen?

The offline metric may not track the business metric, training-serving skew may change features in production, logged data may carry position or selection bias from the old model, or the gain falls in segments that matter little. Check feature parity on live traffic first.

What a strong answer shows: They treat the online test as the truth and debug the offline setup.

How do you tune hyperparameters on a limited GPU budget?

Use Bayesian optimization or a tool like Optuna with early stopping of weak trials (successive halving or ASHA). Tune on a subset, prioritize learning rate, batch size and regularization, and fix seeds to see variance between runs.

What a strong answer shows: They spend budget on the parameters that matter and know run-to-run noise.

Describe a two-stage recommendation architecture.

A retrieval stage narrows millions of items to hundreds using cheap methods such as two-tower embeddings with ANN search, popularity and rules. A ranking stage scores those candidates with a richer model using cross features. A final re-ranking layer applies business rules and diversity.

What a strong answer shows: They explain why stages exist (latency and cost) and how each is evaluated.

How do you evaluate a ranking model beyond a single number?

Use ranking metrics such as NDCG@k and recall@k, then slice by user segment, item popularity and new versus returning users. Check calibration if scores feed downstream decisions, and look at examples of the worst failures.

What a strong answer shows: They look for regressions hidden inside a better average.

MLOps and serving

What is training-serving skew, and how does a feature store help?

Skew happens when features are computed differently in training (batch SQL) and serving (application code), or with different freshness. A feature store such as Feast or a managed equivalent defines features once, materializes them to an offline store for training and an online store for low-latency lookup and supports point-in-time correct training sets.

What a strong answer shows: They know the store solves consistency, not modeling, and add skew monitoring.

Build a point-in-time correct training set from labels and a feature history table.

Each label should only see feature values that existed before the event.

import pandas as pd

train = pd.merge_asof(
labels.sort_values("event_ts"),
features.sort_values("feature_ts"),
left_on="event_ts", right_on="feature_ts",
by="user_id", direction="backward",
)

What a strong answer shows: They spot that a plain join on user ID leaks future values.

How do you detect drift when labels arrive weeks later?

Monitor input distributions and prediction distributions with metrics such as population stability index, track missing-value rates and watch business proxies. Data drift (inputs change) can be measured immediately; concept drift (the relationship changes) needs labels or proxies.

import numpy as np

def psi(expected, actual, bins=10):
edges = np.quantile(expected, np.linspace(0, 1, bins + 1))
edges[0], edges[-1] = -np.inf, np.inf
e = np.clip(np.histogram(expected, edges)[0] / len(expected), 1e-6, None)
a = np.clip(np.histogram(actual, edges)[0] / len(actual), 1e-6, None)
return float(np.sum((a - e) * np.log(a / e)))

What a strong answer shows: They distinguish data drift from concept drift and set alert thresholds tied to action.

What does it take to reproduce a model trained six months ago?

The exact code commit, data snapshot or table version, feature definitions, config and hyperparameters, environment (container image, library versions) and random seeds. A registry such as MLflow ties these to the model artifact. GPU nondeterminism means "reproducible" often means within tolerance.

What a strong answer shows: They version data, not only code.

How do you roll out a new model safely?

Run it in shadow mode on live traffic to compare outputs and latency, then canary to a small share, then A/B test against the current model with guardrail metrics. Keep the previous model deployable for instant rollback.

What a strong answer shows: A staged rollout with clear promotion and rollback criteria.

A ranking model must respond in under 50 ms at p99. How do you meet the budget?

Break down the budget: feature fetch, inference and network. Precompute and cache features, batch requests on the server, export to an optimized runtime such as ONNX Runtime or TensorRT, quantize or distill to a smaller model and cut candidates earlier in retrieval. Measure tail latency under realistic load.

What a strong answer shows: They profile where the time goes before shrinking the model.

When do you choose batch inference over online inference?

Batch works when predictions can be computed ahead of time for a known set of entities, such as daily churn scores. Online is needed when inputs only exist at request time, like a transaction being scored. Batch is cheaper and simpler to monitor.

What a strong answer shows: They default to the simpler option unless freshness demands otherwise.

Senior and architecture

Design a real-time fraud scoring system for card transactions.

A streaming layer computes velocity features (transactions per card in the last minutes), an online store serves them, and a model scores each transaction within the authorization latency budget, with rules as a fallback. Labels arrive late through chargebacks, so training uses label-delay windows and analyst reviews add faster feedback. Monitor score distributions, approval rates and false positive impact on customers.

What a strong answer shows: They handle label delay, fallbacks when the model is down and the cost of false positives.

How do you decide when to retrain a model?

Combine a schedule tied to how fast the domain changes with triggers from drift or performance monitoring. Every retrain goes through the same evaluation gates as a new model, comparing against the current champion on recent data.

What a strong answer shows: They automate retraining with gates, not blind replacement.

How do you avoid feedback loops in a recommender trained on its own logs?

The model only sees outcomes for items it chose to show. Reserve some traffic for exploration, log propensities so you can correct with inverse propensity weighting and evaluate off-policy before shipping.

What a strong answer shows: They see the selection bias in logged data.

How do you check a credit or hiring model for unfair outcomes?

Evaluate error rates and approval rates across relevant groups, choose fairness criteria with legal and policy input, and document limitations. Removing a protected attribute is not enough, since proxies like ZIP code can carry it.

What a strong answer shows: They know fairness metrics can conflict and involve the right stakeholders.

When is an internal ML platform worth building?

When several teams repeat the same work: feature pipelines, training jobs, deployment and monitoring. Before that, managed services and simple conventions are cheaper. Build the pieces that hurt most, often feature consistency and deployment.

What a strong answer shows: They tie platform investment to measured pain.

Training costs on GPUs are growing fast. Where do you look?

Utilization first, since idle GPUs are the most common waste. Then use mixed precision, right-size instances, run preemptible capacity with frequent checkpoints, stop failing runs early and cut duplicate experiments with better tracking.

What a strong answer shows: They treat compute as a budget with owners.

Red flags to watch for

A practical exercise

Run a 90-minute pairing session on a prepared repo. It contains a PyTorch tabular model with a leaky feature, a random split on time-ordered data and an inference script that computes one feature differently from training. Ask the candidate to find the problems, fix them, retrain and add a simple drift check for the serving inputs.

Hire senior machine learning engineers vetted with these questions

Ryz introduces senior machine learning engineers who have trained, deployed and monitored models in production, from fraud scoring to recommendations. They are the top 1% of the candidates we interview, they work on your team, repos and on-call, and they keep hours within ±1h of US time zones. Read about our vetting process, or adapt our machine learning engineer job description for your role.

FAQ

How is an ML engineer interview different from a data scientist interview?

Data scientist interviews center on statistics, experiments and analysis. ML engineer interviews center on building and running models: training code, evaluation pipelines, serving, latency and monitoring. Expect stronger software engineering from an ML engineer.

Should I ask ML engineers to derive algorithms on a whiteboard?

Rarely. A quick check that they understand backpropagation or gradient boosting is fine, but debugging a training run or designing a serving path predicts job performance far better.

How do I test ML skills remotely?

Use a shared repo with a small dataset that runs on a laptop CPU, so hardware does not decide the outcome. Pair live on debugging, then discuss how the solution would change at production scale.

Questions we didn't answer? Email info@ryzlabs.com.

Explore Ryz Labs

Staff augmentationDedicated development teamsAI pod teamsForward deployed engineersNearshore software developmentAI engineering teamsHire engineers by roleRyz Labs vs competitorsAlternatives guidesBuyer guidesCase studiesHow we vet engineers
Ryz Labs

Senior engineers in your time zone. AI pod teams that ship.

Tell us who you need. You'll get a scoped plan, a price and the names of the people who would do the work.

Start a conversation →