These 27 machine learning engineer interview questions focus on getting models trained, evaluated, served and kept healthy in production: PyTorch training loops and their failure modes, distributed training, offline versus online evaluation, feature stores, deployment, latency budgets and drift. They are written for senior ML engineers rather than researchers or analysts. At Ryz, ML engineers are evaluated along the same lines: recruiters source people who have owned models in production, candidates complete structured NTRVSTA AI interviews on training, serving and system design, and recruiters review every candidate before and after. AI scores are advisory, and people make the hiring decisions.
Decide which half of the role matters more. Some ML engineers spend their time on modeling and training; others build the platform and serving path around models. Weight the intermediate section for the first and the MLOps section for the second, and use the senior section for anyone who will own a system end to end.
Ask for specifics from production: the last model they rolled back and why, how they knew it was degrading, what the latency budget was. Engineers who have shipped answer with numbers and incidents. Finish with a practical exercise, since a working training and evaluation script reveals habits that conversation hides.
model.train()
for x, y in train_loader:
x, y = x.to(device), y.to(device)
optimizer.zero_grad(set_to_none=True)
loss = loss_fn(model(x), y)
loss.backward()
optimizer.step()
model.eval()
with torch.no_grad():
val_loss = sum(loss_fn(model(x.to(device)), y.to(device)).item()
for x, y in val_loader) / len(val_loader)
Common bugs: forgetting model.eval(), so dropout and batch norm behave as in training; computing validation without no_grad, which wastes memory; and accumulating loss tensors instead of .item(), which keeps graphs alive.
What a strong answer shows: They write it fluently and explain train versus eval mode without being asked.
The model is overfitting. Keep the checkpoint with the best validation metric (early stopping), then add regularization: weight decay, dropout, data augmentation or a smaller model. Also check whether the validation set matches production data, because a distribution gap looks similar.
What a strong answer shows: They checkpoint on validation and consider data problems, not only regularization knobs.
Trees handle heterogeneous feature scales, missing values and sharp thresholds well with little tuning, and train fast on CPUs. Neural networks pay off on tabular data mainly when you have very large datasets, high-cardinality categorical features that benefit from embeddings, or need to combine tabular inputs with text or images.
What a strong answer shows: They start with a strong baseline like LightGBM or XGBoost before reaching for deep learning.
It sets the step size of each update. Too high and loss diverges or oscillates; too low and training crawls or stalls. A common recipe is AdamW with linear warmup followed by cosine decay. A short learning-rate range test helps find a sensible peak.
What a strong answer shows: They explain why warmup stabilizes early training, especially for transformers.
An embedding is a dense vector learned so that similar items are close together. In recommendations, user and item embeddings from a two-tower model enable fast candidate retrieval with approximate nearest neighbor search. In fraud, merchant embeddings capture behavior that one-hot encoding cannot.
What a strong answer shows: They connect embeddings to a concrete retrieval or feature use case.
Batch norm normalizes each feature across the batch, so it depends on batch statistics and behaves differently at inference. Layer norm normalizes across features within one example, which works with variable sequence lengths and small or per-token batches. Most modern transformers also place the norm before attention and MLP blocks for stability.
What a strong answer shows: They tie the choice to batch-size dependence and inference behavior.
Split by time: train on the past, validate on the following period and test on the most recent. Group by entity if the same user can appear in both sets. Touch the test set once, at the end, or it stops being a test.
What a strong answer shows: They mirror production conditions in the split.
Check for a learning rate spike, exploding gradients, a bad batch (divide by zero, log of zero) and fp16 overflow. Log gradient norms, add clipping, prefer bf16 on hardware that supports it, and if using fp16, make sure a gradient scaler is in place.
scaler = torch.amp.GradScaler("cuda")
with torch.autocast(device_type="cuda", dtype=torch.float16):
loss = loss_fn(model(x), y)
scaler.scale(loss).backward()
scaler.unscale_(optimizer)
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
scaler.step(optimizer)
scaler.update()
What a strong answer shows: They know to unscale before clipping and can reproduce the failing batch from a checkpoint.
Usually the input pipeline. Profile with the PyTorch profiler. Fixes include more DataLoader workers, pin_memory=True, preprocessing offline, faster formats, larger batches and moving augmentation to the GPU. torch.compile can reduce Python overhead once the GPU is actually busy.
What a strong answer shows: They measure before tuning and know the data loader is the usual suspect.
DDP replicates the full model on each GPU and averages gradients; it is simple and fast when the model, gradients and optimizer state fit on one device. FSDP shards parameters, gradients and optimizer state across GPUs, which is needed for models that do not fit, at the cost of more communication.
What a strong answer shows: They reason about memory per GPU, including optimizer state.
The offline metric may not track the business metric, training-serving skew may change features in production, logged data may carry position or selection bias from the old model, or the gain falls in segments that matter little. Check feature parity on live traffic first.
What a strong answer shows: They treat the online test as the truth and debug the offline setup.
Use Bayesian optimization or a tool like Optuna with early stopping of weak trials (successive halving or ASHA). Tune on a subset, prioritize learning rate, batch size and regularization, and fix seeds to see variance between runs.
What a strong answer shows: They spend budget on the parameters that matter and know run-to-run noise.
A retrieval stage narrows millions of items to hundreds using cheap methods such as two-tower embeddings with ANN search, popularity and rules. A ranking stage scores those candidates with a richer model using cross features. A final re-ranking layer applies business rules and diversity.
What a strong answer shows: They explain why stages exist (latency and cost) and how each is evaluated.
Use ranking metrics such as NDCG@k and recall@k, then slice by user segment, item popularity and new versus returning users. Check calibration if scores feed downstream decisions, and look at examples of the worst failures.
What a strong answer shows: They look for regressions hidden inside a better average.
Skew happens when features are computed differently in training (batch SQL) and serving (application code), or with different freshness. A feature store such as Feast or a managed equivalent defines features once, materializes them to an offline store for training and an online store for low-latency lookup and supports point-in-time correct training sets.
What a strong answer shows: They know the store solves consistency, not modeling, and add skew monitoring.
Each label should only see feature values that existed before the event.
import pandas as pd
train = pd.merge_asof(
labels.sort_values("event_ts"),
features.sort_values("feature_ts"),
left_on="event_ts", right_on="feature_ts",
by="user_id", direction="backward",
)
What a strong answer shows: They spot that a plain join on user ID leaks future values.
Monitor input distributions and prediction distributions with metrics such as population stability index, track missing-value rates and watch business proxies. Data drift (inputs change) can be measured immediately; concept drift (the relationship changes) needs labels or proxies.
import numpy as np
def psi(expected, actual, bins=10):
edges = np.quantile(expected, np.linspace(0, 1, bins + 1))
edges[0], edges[-1] = -np.inf, np.inf
e = np.clip(np.histogram(expected, edges)[0] / len(expected), 1e-6, None)
a = np.clip(np.histogram(actual, edges)[0] / len(actual), 1e-6, None)
return float(np.sum((a - e) * np.log(a / e)))
What a strong answer shows: They distinguish data drift from concept drift and set alert thresholds tied to action.
The exact code commit, data snapshot or table version, feature definitions, config and hyperparameters, environment (container image, library versions) and random seeds. A registry such as MLflow ties these to the model artifact. GPU nondeterminism means "reproducible" often means within tolerance.
What a strong answer shows: They version data, not only code.
Run it in shadow mode on live traffic to compare outputs and latency, then canary to a small share, then A/B test against the current model with guardrail metrics. Keep the previous model deployable for instant rollback.
What a strong answer shows: A staged rollout with clear promotion and rollback criteria.
Break down the budget: feature fetch, inference and network. Precompute and cache features, batch requests on the server, export to an optimized runtime such as ONNX Runtime or TensorRT, quantize or distill to a smaller model and cut candidates earlier in retrieval. Measure tail latency under realistic load.
What a strong answer shows: They profile where the time goes before shrinking the model.
Batch works when predictions can be computed ahead of time for a known set of entities, such as daily churn scores. Online is needed when inputs only exist at request time, like a transaction being scored. Batch is cheaper and simpler to monitor.
What a strong answer shows: They default to the simpler option unless freshness demands otherwise.
A streaming layer computes velocity features (transactions per card in the last minutes), an online store serves them, and a model scores each transaction within the authorization latency budget, with rules as a fallback. Labels arrive late through chargebacks, so training uses label-delay windows and analyst reviews add faster feedback. Monitor score distributions, approval rates and false positive impact on customers.
What a strong answer shows: They handle label delay, fallbacks when the model is down and the cost of false positives.
Combine a schedule tied to how fast the domain changes with triggers from drift or performance monitoring. Every retrain goes through the same evaluation gates as a new model, comparing against the current champion on recent data.
What a strong answer shows: They automate retraining with gates, not blind replacement.
The model only sees outcomes for items it chose to show. Reserve some traffic for exploration, log propensities so you can correct with inverse propensity weighting and evaluate off-policy before shipping.
What a strong answer shows: They see the selection bias in logged data.
Evaluate error rates and approval rates across relevant groups, choose fairness criteria with legal and policy input, and document limitations. Removing a protected attribute is not enough, since proxies like ZIP code can carry it.
What a strong answer shows: They know fairness metrics can conflict and involve the right stakeholders.
When several teams repeat the same work: feature pipelines, training jobs, deployment and monitoring. Before that, managed services and simple conventions are cheaper. Build the pieces that hurt most, often feature consistency and deployment.
What a strong answer shows: They tie platform investment to measured pain.
Utilization first, since idle GPUs are the most common waste. Then use mixed precision, right-size instances, run preemptible capacity with frequent checkpoints, stop failing runs early and cut duplicate experiments with better tracking.
What a strong answer shows: They treat compute as a budget with owners.
model.eval(), validation splits by time or point-in-time features.Run a 90-minute pairing session on a prepared repo. It contains a PyTorch tabular model with a leaky feature, a random split on time-ordered data and an inference script that computes one feature differently from training. Ask the candidate to find the problems, fix them, retrain and add a simple drift check for the serving inputs.
Ryz introduces senior machine learning engineers who have trained, deployed and monitored models in production, from fraud scoring to recommendations. They are the top 1% of the candidates we interview, they work on your team, repos and on-call, and they keep hours within ±1h of US time zones. Read about our vetting process, or adapt our machine learning engineer job description for your role.
Data scientist interviews center on statistics, experiments and analysis. ML engineer interviews center on building and running models: training code, evaluation pipelines, serving, latency and monitoring. Expect stronger software engineering from an ML engineer.
Rarely. A quick check that they understand backpropagation or gradient boosting is fine, but debugging a training run or designing a serving path predicts job performance far better.
Use a shared repo with a small dataset that runs on a laptop CPU, so hardware does not decide the outcome. Pair live on debugging, then discuss how the solution would change at production scale.
Questions we didn't answer? Email info@ryzlabs.com.