These 26 data scientist interview questions test the skills that make a senior data scientist useful to a business: sound statistics, well-designed experiments, causal reasoning when you cannot randomize, models that hold up outside the notebook and the ability to explain an uncertain result to someone who has to act on it. At Ryz, senior data scientists go through the same kind of evaluation. Recruiters source people with real analytical ownership, candidates complete structured NTRVSTA AI interviews on statistics and experimentation, and recruiters review each candidate before and after. AI scores are advisory; people make the decisions.
Match the questions to the role. A product analytics data scientist should spend most of the interview in the experimentation section. A data scientist who builds predictive models should get more of the intermediate modeling questions. Senior candidates of either kind should handle the senior section, because it is about framing problems and influencing decisions.
Do not reward formula recall. Ask the candidate to explain a concept as they would to a product manager, then push on an edge case. Add a short analysis exercise with a messy dataset, since judgment about data quality rarely shows up in conversation alone.
It means that if the null hypothesis were true, data at least this extreme would appear about 3% of the time. It does not mean the treatment works with probability 0.97, and it says nothing about the size or business value of the effect.
What a strong answer shows: They correct the common misreading unprompted and pair p-values with effect sizes and confidence intervals.
A Type I error is a false positive, controlled by alpha. A Type II error is a false negative, and power is one minus its probability. For a fixed alpha, more samples raise power, so smaller true effects become detectable. Underpowered tests produce noise and, when they do hit significance, exaggerated effect estimates.
What a strong answer shows: They mention that significant results from underpowered tests overstate the effect, not only that they are rare.
Simple models miss real structure (high bias); flexible models fit noise in the training set (high variance). In practice you manage it with regularization, more data, cross-validation and by comparing training and validation error. A large gap means variance; both errors high means bias.
What a strong answer shows: They diagnose from learning curves rather than reciting a definition.
Not necessarily: predicting "no churn" for everyone also scores 96%. Look at precision and recall for the churn class, the precision-recall curve and average precision, and ask what the retention team can act on. If they can call 500 customers a week, precision in the top 500 matters most.
What a strong answer shows: They choose a metric from how the prediction will be used.
A trend in aggregated data reverses within every subgroup. A new onboarding flow may lower overall conversion while raising it on both mobile and desktop, because it launched during a period when traffic shifted toward lower-converting mobile users. The mix changed, not the flow's effect.
What a strong answer shows: They check segment mix before trusting a top-line comparison.
First ask why it is missing. Missing completely at random is rare; missing at random can be modeled with other variables; missing not at random (income left blank by high earners) biases any simple fix. Options include dropping, imputation with an indicator column, model-based imputation or models that handle missing values natively.
What a strong answer shows: They treat missingness as information and check whether it correlates with the target.
No. Engaged users are more likely to find feature X and to stay, so engagement confounds the comparison. You need an experiment, or a quasi-experimental design with a credible source of variation, before claiming the feature drives retention.
What a strong answer shows: They name the likely confounder and propose how to test the claim.
Leakage. Common causes are features computed with information from after the prediction time, random splits on time-dependent data and duplicate entities across train and test. Rebuild features as of the prediction timestamp and validate with a time-based split.
from sklearn.model_selection import TimeSeriesSplit, cross_val_score
cv = TimeSeriesSplit(n_splits=5, gap=7) # rows sorted by time; 7-row gap
scores = cross_val_score(model, X, y, cv=cv, scoring="average_precision")
What a strong answer shows: They ask "what would we have known at prediction time?" for every feature.
A model can rank well and still output probabilities that are systematically too high or low. If prices depend on expected loss, miscalibration costs money directly. Check with a reliability curve and the Brier score, then calibrate on held-out data.
from sklearn.calibration import CalibratedClassifierCV, calibration_curve
calibrated = CalibratedClassifierCV(model, method="isotonic", cv=5).fit(X_train, y_train)
p = calibrated.predict_proba(X_test)[:, 1]
prob_true, prob_pred = calibration_curve(y_test, p, n_bins=10)
What a strong answer shows: They separate ranking quality from probability quality and know isotonic needs more data than sigmoid.
From costs, not 0.5. Estimate the cost of a false positive and a false negative, then pick the threshold that minimizes expected cost on validation data, or the one that fits a capacity limit.
import numpy as np
thresholds = np.linspace(0.05, 0.95, 91)
costs = [fp_cost * ((p >= t) & (y == 0)).sum() + fn_cost * ((p < t) & (y == 1)).sum()
for t in thresholds]
best = thresholds[int(np.argmin(costs))]
What a strong answer shows: They tie the threshold to a business decision and revisit it when costs change.
with first_seen as (
select user_id, min(event_date) as cohort_date
from events
group by user_id
)
select f.cohort_date,
count(distinct f.user_id) as cohort_size,
count(distinct case when e.event_date between f.cohort_date + 1
and f.cohort_date + 7
then e.user_id end) as retained_d7
from first_seen f
left join events e on e.user_id = f.user_id
group by f.cohort_date
order by f.cohort_date;
Recent cohorts have not had seven days yet, so exclude them or mark them incomplete.
What a strong answer shows: They define retention precisely and handle immature cohorts.
L2 (ridge) shrinks all coefficients toward zero and handles correlated features by spreading weight across them. L1 (lasso) can push coefficients exactly to zero, which gives feature selection but picks one of a correlated group somewhat arbitrarily. Elastic net combines both. Standardize features first, since both penalties depend on scale.
What a strong answer shows: They mention scaling and the instability of lasso selection with correlated features.
Start with the right metric and a threshold chosen from costs. Then try class weights, which change the loss without discarding data. Resampling such as undersampling or SMOTE can help some models but distorts predicted probabilities, so recalibrate afterward.
What a strong answer shows: They know resampling breaks calibration.
SHAP values and permutation importance describe what the model relies on, not what causes the outcome. With correlated features, importance gets split or shuffled between them. Present them as model behavior, check stability across retrains and avoid causal language.
What a strong answer shows: They keep explanations of a model separate from claims about the world.
State the hypothesis, the randomization unit (user, not session), one primary metric (completed purchases per user), guardrails (refund rate, latency, support contacts) and the minimum detectable effect worth acting on. Compute sample size and duration up front, covering at least one full weekly cycle.
from statsmodels.stats.power import NormalIndPower
from statsmodels.stats.proportion import proportion_effectsize
effect = proportion_effectsize(0.105, 0.100) # 10.0% baseline, +0.5pt MDE
n_per_arm = NormalIndPower().solve_power(effect_size=effect, alpha=0.05, power=0.8)
What a strong answer shows: They decide the metric, MDE and stopping rule before launch.
Repeated looks at a fixed-horizon test inflate the false positive rate well above the nominal 5%. Either commit to a fixed sample size or use a method designed for monitoring, such as group sequential designs with alpha spending or always-valid sequential tests.
What a strong answer shows: They offer a way to monitor safely, not just a refusal.
Not yet. That is a sample ratio mismatch, and with this many users it is far beyond chance. It usually means a bug in assignment, redirects losing users or logging that differs by variant. Results are untrustworthy until the cause is found.
from scipy.stats import chisquare
stat, p = chisquare([1_016_000, 984_000], f_exp=[1_000_000, 1_000_000])
What a strong answer shows: They check SRM before looking at metrics and know where it comes from.
CUPED reduces variance by adjusting the outcome with a pre-experiment covariate, usually the same metric before the test. The more correlated the covariate, the more variance falls, so tests reach the same power with fewer users. It does little for new users with no history.
import numpy as np
theta = np.cov(pre, post)[0, 1] / np.var(pre, ddof=1)
post_adj = post - theta * (pre - pre.mean())
What a strong answer shows: They know the covariate must be unaffected by treatment.
Sessions from the same user are correlated, so treating them as independent understates variance and produces false positives. Use the delta method for ratio metrics or a bootstrap that resamples users, not sessions.
What a strong answer shows: They match the analysis unit to the randomization unit.
Use difference-in-differences against comparable regions, after checking that pre-launch trends were parallel. If no single region is comparable, build a synthetic control from a weighted mix. Run placebo tests on dates and regions where nothing launched.
What a strong answer shows: They state the assumptions and test them rather than presenting one number.
Interference breaks the independence assumption: if treated riders book more drivers, control riders see fewer. Use cluster randomization by city or region, or switchback designs that alternate treatment over time windows, and accept wider confidence intervals.
What a strong answer shows: They spot interference before it biases the result.
Ask what action follows a prediction. If the plan is a retention offer, the useful question is who will stay because of the offer, which is uplift modeling. A propensity model will target customers who would have left anyway or stayed anyway. Agree on a test that measures retained revenue from the intervention.
What a strong answer shows: They push from prediction to decision and define success in business terms.
If the segment was not pre-registered, treat it as a hypothesis. Slicing by many segments guarantees some false positives. Check whether the effect is plausible and large, then run a confirmatory test on new users.
What a strong answer shows: They protect decision quality without blocking the team.
Plot the treatment effect over time to see whether it decays. Keep a long-running holdout for major launches, and validate short-term proxy metrics against long-term outcomes before optimizing for them.
What a strong answer shows: They question whether the proxy metric is actually predictive.
Write one definition with an owner, implement it once in a governed semantic or metrics layer and point dashboards at it. Publish the definition and its known caveats so arguments shift from whose number is right to whether the definition fits the question.
What a strong answer shows: They treat metric definitions as shared infrastructure.
Lead with the decision and what the data supports: "We can rule out an effect larger than +1%, so this will not hit the 5% goal." Show the confidence interval, the cost of continuing and a recommendation. Inconclusive still carries information.
What a strong answer shows: They turn uncertainty into a usable recommendation.
Give a three-hour take-home with raw event data from a completed A/B test on a pricing page: assignment logs, sessions and purchases, with a planted sample ratio issue in one platform and some duplicate events. Ask for a short memo to a product leader recommending ship, iterate or stop, plus the notebook behind it.
Ryz introduces senior data scientists who have designed experiments, built causal analyses and shipped models that changed product decisions. They are the top 1% of the candidates we interview, they work on your team, joining your standups and analytics reviews, and they keep hours within ±1h of US time zones. See how candidates are screened in our vetting process, or start from our data scientist job description.
Yes, but test the coding they will do: SQL, pandas or Polars and a modeling library. A short live exercise on cleaning and summarizing data is more predictive than algorithm puzzles.
Mid-level candidates answer the question asked well. Senior candidates question the framing, ask what decision the analysis supports and explain the limits of their result without being prompted.
A take-home with a written memo shows analysis quality and communication. A live case shows how they reason under questioning. If you can only do one, use a short take-home and review it together live.
Questions we didn't answer? Email info@ryzlabs.com.