Interview questions

Data scientist interview questions for senior hires (2026)

Statistics, experiment design, causal inference and modeling questions that show whether a data scientist can change a decision, not just fit a model.

These 26 data scientist interview questions test the skills that make a senior data scientist useful to a business: sound statistics, well-designed experiments, causal reasoning when you cannot randomize, models that hold up outside the notebook and the ability to explain an uncertain result to someone who has to act on it. At Ryz, senior data scientists go through the same kind of evaluation. Recruiters source people with real analytical ownership, candidates complete structured NTRVSTA AI interviews on statistics and experimentation, and recruiters review each candidate before and after. AI scores are advisory; people make the decisions.

How to use these questions

Match the questions to the role. A product analytics data scientist should spend most of the interview in the experimentation section. A data scientist who builds predictive models should get more of the intermediate modeling questions. Senior candidates of either kind should handle the senior section, because it is about framing problems and influencing decisions.

Do not reward formula recall. Ask the candidate to explain a concept as they would to a product manager, then push on an edge case. Add a short analysis exercise with a messy dataset, since judgment about data quality rarely shows up in conversation alone.

Fundamentals

What does a p-value of 0.03 mean, and what does it not mean?

It means that if the null hypothesis were true, data at least this extreme would appear about 3% of the time. It does not mean the treatment works with probability 0.97, and it says nothing about the size or business value of the effect.

What a strong answer shows: They correct the common misreading unprompted and pair p-values with effect sizes and confidence intervals.

Explain Type I and Type II errors, and how sample size affects each.

A Type I error is a false positive, controlled by alpha. A Type II error is a false negative, and power is one minus its probability. For a fixed alpha, more samples raise power, so smaller true effects become detectable. Underpowered tests produce noise and, when they do hit significance, exaggerated effect estimates.

What a strong answer shows: They mention that significant results from underpowered tests overstate the effect, not only that they are rare.

What is the bias-variance trade-off, in practical terms?

Simple models miss real structure (high bias); flexible models fit noise in the training set (high variance). In practice you manage it with regularization, more data, cross-validation and by comparing training and validation error. A large gap means variance; both errors high means bias.

What a strong answer shows: They diagnose from learning curves rather than reciting a definition.

A churn model has 96% accuracy on data where 4% of customers churn. Is it good?

Not necessarily: predicting "no churn" for everyone also scores 96%. Look at precision and recall for the churn class, the precision-recall curve and average precision, and ask what the retention team can act on. If they can call 500 customers a week, precision in the top 500 matters most.

What a strong answer shows: They choose a metric from how the prediction will be used.

Explain Simpson's paradox with an example you might see at work.

A trend in aggregated data reverses within every subgroup. A new onboarding flow may lower overall conversion while raising it on both mobile and desktop, because it launched during a period when traffic shifted toward lower-converting mobile users. The mix changed, not the flow's effect.

What a strong answer shows: They check segment mix before trusting a top-line comparison.

How do you handle missing data?

First ask why it is missing. Missing completely at random is rare; missing at random can be modeled with other variables; missing not at random (income left blank by high earners) biases any simple fix. Options include dropping, imputation with an indicator column, model-based imputation or models that handle missing values natively.

What a strong answer shows: They treat missingness as information and check whether it correlates with the target.

Your analysis shows users who use feature X retain twice as well. Can you say feature X causes retention?

No. Engaged users are more likely to find feature X and to stay, so engagement confounds the comparison. You need an experiment, or a quasi-experimental design with a credible source of variation, before claiming the feature drives retention.

What a strong answer shows: They name the likely confounder and propose how to test the claim.

Intermediate

A model scored 0.97 AUC in validation and performs barely better than chance in production. What do you check first?

Leakage. Common causes are features computed with information from after the prediction time, random splits on time-dependent data and duplicate entities across train and test. Rebuild features as of the prediction timestamp and validate with a time-based split.

from sklearn.model_selection import TimeSeriesSplit, cross_val_score

cv = TimeSeriesSplit(n_splits=5, gap=7) # rows sorted by time; 7-row gap
scores = cross_val_score(model, X, y, cv=cv, scoring="average_precision")

What a strong answer shows: They ask "what would we have known at prediction time?" for every feature.

Your model's probabilities will be used to set prices. Why does calibration matter, and how do you check it?

A model can rank well and still output probabilities that are systematically too high or low. If prices depend on expected loss, miscalibration costs money directly. Check with a reliability curve and the Brier score, then calibrate on held-out data.

from sklearn.calibration import CalibratedClassifierCV, calibration_curve

calibrated = CalibratedClassifierCV(model, method="isotonic", cv=5).fit(X_train, y_train)
p = calibrated.predict_proba(X_test)[:, 1]
prob_true, prob_pred = calibration_curve(y_test, p, n_bins=10)

What a strong answer shows: They separate ranking quality from probability quality and know isotonic needs more data than sigmoid.

How do you choose a classification threshold?

From costs, not 0.5. Estimate the cost of a false positive and a false negative, then pick the threshold that minimizes expected cost on validation data, or the one that fits a capacity limit.

import numpy as np

thresholds = np.linspace(0.05, 0.95, 91)
costs = [fp_cost * ((p >= t) & (y == 0)).sum() + fn_cost * ((p < t) & (y == 1)).sum()
for t in thresholds]
best = thresholds[int(np.argmin(costs))]

What a strong answer shows: They tie the threshold to a business decision and revisit it when costs change.

Write SQL for day-7 retention by signup cohort.

with first_seen as (
select user_id, min(event_date) as cohort_date
from events
group by user_id
)
select f.cohort_date,
count(distinct f.user_id) as cohort_size,
count(distinct case when e.event_date between f.cohort_date + 1
and f.cohort_date + 7
then e.user_id end) as retained_d7
from first_seen f
left join events e on e.user_id = f.user_id
group by f.cohort_date
order by f.cohort_date;

Recent cohorts have not had seven days yet, so exclude them or mark them incomplete.

What a strong answer shows: They define retention precisely and handle immature cohorts.

Compare L1 and L2 regularization. When would you pick each?

L2 (ridge) shrinks all coefficients toward zero and handles correlated features by spreading weight across them. L1 (lasso) can push coefficients exactly to zero, which gives feature selection but picks one of a correlated group somewhat arbitrarily. Elastic net combines both. Standardize features first, since both penalties depend on scale.

What a strong answer shows: They mention scaling and the instability of lasso selection with correlated features.

How do you deal with a heavily imbalanced target?

Start with the right metric and a threshold chosen from costs. Then try class weights, which change the loss without discarding data. Resampling such as undersampling or SMOTE can help some models but distorts predicted probabilities, so recalibrate afterward.

What a strong answer shows: They know resampling breaks calibration.

A stakeholder asks which features "drive" the model. How do you answer honestly?

SHAP values and permutation importance describe what the model relies on, not what causes the outcome. With correlated features, importance gets split or shuffled between them. Present them as model behavior, check stability across retrains and avoid causal language.

What a strong answer shows: They keep explanations of a model separate from claims about the world.

Experimentation and causal inference

Design an A/B test for a new checkout flow.

State the hypothesis, the randomization unit (user, not session), one primary metric (completed purchases per user), guardrails (refund rate, latency, support contacts) and the minimum detectable effect worth acting on. Compute sample size and duration up front, covering at least one full weekly cycle.

from statsmodels.stats.power import NormalIndPower
from statsmodels.stats.proportion import proportion_effectsize

effect = proportion_effectsize(0.105, 0.100) # 10.0% baseline, +0.5pt MDE
n_per_arm = NormalIndPower().solve_power(effect_size=effect, alpha=0.05, power=0.8)

What a strong answer shows: They decide the metric, MDE and stopping rule before launch.

A product manager checks the dashboard daily and wants to stop as soon as the result is significant. What is the problem?

Repeated looks at a fixed-horizon test inflate the false positive rate well above the nominal 5%. Either commit to a fixed sample size or use a method designed for monitoring, such as group sequential designs with alpha spending or always-valid sequential tests.

What a strong answer shows: They offer a way to monitor safely, not just a refusal.

Your 50/50 test shows 1,016,000 users in control and 984,000 in treatment. Should you analyze it?

Not yet. That is a sample ratio mismatch, and with this many users it is far beyond chance. It usually means a bug in assignment, redirects losing users or logging that differs by variant. Results are untrustworthy until the cause is found.

from scipy.stats import chisquare

stat, p = chisquare([1_016_000, 984_000], f_exp=[1_000_000, 1_000_000])

What a strong answer shows: They check SRM before looking at metrics and know where it comes from.

What is CUPED and when does it help?

CUPED reduces variance by adjusting the outcome with a pre-experiment covariate, usually the same metric before the test. The more correlated the covariate, the more variance falls, so tests reach the same power with fewer users. It does little for new users with no history.

import numpy as np

theta = np.cov(pre, post)[0, 1] / np.var(pre, ddof=1)
post_adj = post - theta * (pre - pre.mean())

What a strong answer shows: They know the covariate must be unaffected by treatment.

You randomize by user but report revenue per session. What goes wrong with a standard t-test?

Sessions from the same user are correlated, so treating them as independent understates variance and produces false positives. Use the delta method for ratio metrics or a bootstrap that resamples users, not sessions.

What a strong answer shows: They match the analysis unit to the randomization unit.

A feature launched to one region with no holdout. How do you estimate its impact?

Use difference-in-differences against comparable regions, after checking that pre-launch trends were parallel. If no single region is comparable, build a synthetic control from a weighted mix. Run placebo tests on dates and regions where nothing launched.

What a strong answer shows: They state the assumptions and test them rather than presenting one number.

How do you run experiments in a two-sided marketplace where treatment users affect control users?

Interference breaks the independence assumption: if treated riders book more drivers, control riders see fewer. Use cluster randomization by city or region, or switchback designs that alternate treatment over time windows, and accept wider confidence intervals.

What a strong answer shows: They spot interference before it biases the result.

Senior and architecture

Leadership asks for a churn prediction model. How do you frame the project?

Ask what action follows a prediction. If the plan is a retention offer, the useful question is who will stay because of the offer, which is uplift modeling. A propensity model will target customers who would have left anyway or stayed anyway. Agree on a test that measures retained revenue from the intervention.

What a strong answer shows: They push from prediction to decision and define success in business terms.

A test is flat overall but positive for new users. Leadership wants to ship it to new users only. What do you advise?

If the segment was not pre-registered, treat it as a hypothesis. Slicing by many segments guarantees some false positives. Check whether the effect is plausible and large, then run a confirmatory test on new users.

What a strong answer shows: They protect decision quality without blocking the team.

How do you handle novelty effects and metrics that move slowly, like long-term retention?

Plot the treatment effect over time to see whether it decays. Keep a long-running holdout for major launches, and validate short-term proxy metrics against long-term outcomes before optimizing for them.

What a strong answer shows: They question whether the proxy metric is actually predictive.

Different teams report different numbers for "active users". How do you fix it?

Write one definition with an owner, implement it once in a governed semantic or metrics layer and point dashboards at it. Publish the definition and its known caveats so arguments shift from whose number is right to whether the definition fits the question.

What a strong answer shows: They treat metric definitions as shared infrastructure.

How do you present an inconclusive experiment to an executive?

Lead with the decision and what the data supports: "We can rule out an effect larger than +1%, so this will not hit the 5% goal." Show the confidence interval, the cost of continuing and a recommendation. Inconclusive still carries information.

What a strong answer shows: They turn uncertainty into a usable recommendation.

Red flags to watch for

A practical exercise

Give a three-hour take-home with raw event data from a completed A/B test on a pricing page: assignment logs, sessions and purchases, with a planted sample ratio issue in one platform and some duplicate events. Ask for a short memo to a product leader recommending ship, iterate or stop, plus the notebook behind it.

Hire senior data scientists vetted with these questions

Ryz introduces senior data scientists who have designed experiments, built causal analyses and shipped models that changed product decisions. They are the top 1% of the candidates we interview, they work on your team, joining your standups and analytics reviews, and they keep hours within ±1h of US time zones. See how candidates are screened in our vetting process, or start from our data scientist job description.

FAQ

Should a data scientist interview include coding?

Yes, but test the coding they will do: SQL, pandas or Polars and a modeling library. A short live exercise on cleaning and summarizing data is more predictive than algorithm puzzles.

How do I tell a senior data scientist from a mid-level one?

Mid-level candidates answer the question asked well. Senior candidates question the framing, ask what decision the analysis supports and explain the limits of their result without being prompted.

Is a take-home or a live case better for data scientists?

A take-home with a written memo shows analysis quality and communication. A live case shows how they reason under questioning. If you can only do one, use a short take-home and review it together live.

Questions we didn't answer? Email info@ryzlabs.com.

Explore Ryz Labs

Staff augmentationDedicated development teamsAI pod teamsForward deployed engineersNearshore software developmentAI engineering teamsHire engineers by roleRyz Labs vs competitorsAlternatives guidesBuyer guidesCase studiesHow we vet engineers
Ryz Labs

Senior engineers in your time zone. AI pod teams that ship.

Tell us who you need. You'll get a scoped plan, a price and the names of the people who would do the work.

Start a conversation →