Interview questions

LLM engineer interview questions for senior hires (2026)

Deep questions on how large language models behave and run: tokens, context, fine-tuning, inference optimization, structured outputs and evaluation.

These 27 LLM engineer interview questions go below the API surface. They cover tokenization and context windows, decoding, structured outputs, prompt caching, the fine-tuning versus retrieval decision, LoRA, KV cache memory, quantization, serving with engines like vLLM, eval harnesses and hallucination measurement. They suit engineers who own model behavior and inference, whether they work with hosted models from providers such as Anthropic and OpenAI or run open-weight models themselves. At Ryz, senior LLM engineers are evaluated on this depth: recruiters source engineers with hands-on model work, candidates complete structured NTRVSTA AI interviews on LLM internals and system design, and recruiters review each candidate before and after. AI scores are advisory; people make the decisions.

How to use these questions

First decide whether the role is mostly hosted-API work or self-hosted inference and fine-tuning. For hosted work, focus on the fundamentals, structured outputs, caching and evaluation questions. For self-hosted work, spend most of the interview in the inference section, where memory math and serving trade-offs quickly separate people who have run models from people who have read about them.

Ask candidates to estimate numbers out loud: tokens per request, KV cache size, GPU count. Exact answers matter less than a sound method. Close with a practical exercise that requires measuring, not just prompting.

Fundamentals

What is a token, and why does tokenization matter in production?

Models read subword tokens produced by a tokenizer specific to the model family. Token counts drive cost, latency and context limits, and they vary by language, formatting and content type: code, numbers and many non-English languages use more tokens per word. Count with the tokenizer for the model you actually call.

import tiktoken

enc = tiktoken.get_encoding("o200k_base")
n_tokens = len(enc.encode(document_text))

What a strong answer shows: They know token counts are model-specific and use them for budgeting, not word counts.

How do temperature and top-p change generation, and why is temperature 0 not fully deterministic?

Temperature rescales the token probability distribution; top-p samples only from the smallest set of tokens whose probabilities add up to p. Temperature 0 approximates greedy decoding, but batching, floating-point non-associativity on GPUs and provider-side changes can still produce different outputs for the same input.

What a strong answer shows: They design for variance with evals rather than assuming reproducibility.

Explain prefill and decode. Why are output tokens slower and usually more expensive than input tokens?

Prefill processes the whole prompt in parallel and builds the KV cache; it is compute-bound. Decode generates one token at a time, each step reading the model weights and cache; it is memory-bandwidth-bound and sequential. That is why time to first token depends on prompt length while total time depends mostly on output length.

What a strong answer shows: They use this to reason about latency, such as shortening outputs before shortening prompts.

Explain self-attention in about a minute.

Each token is projected to query, key and value vectors. A token's output is a weighted sum of all visible tokens' values, with weights from the softmax of its query's dot products with their keys. Multiple heads learn different relationships, and a causal mask stops tokens from seeing the future during generation.

What a strong answer shows: A clear, correct explanation without hand-waving or math theater.

A model supports a very long context window. Why not put every document into every prompt?

Cost and latency grow with input length, and models use information unevenly across long contexts, often missing details in the middle. Long context suits tasks needing a whole document; retrieval is better when only a small slice is relevant or the corpus is large.

What a strong answer shows: They test long-context recall on their own data rather than trusting the advertised limit.

Why do LLMs hallucinate?

They generate plausible continuations, not verified facts. Training rewards fluent answers, knowledge is compressed and imperfect, and models often answer instead of abstaining. Hallucination rises with obscure facts, long-tail entities and prompts that presuppose false information.

What a strong answer shows: They connect causes to mitigations: grounding, abstention options and verification.

What is the difference between a base model and an instruction-tuned model?

A base model is trained to predict the next token on large corpora. An instruction-tuned model is further trained with supervised examples and preference optimization (RLHF or methods like DPO) to follow instructions and match a chat format. Fine-tuning usually starts from the instruct version and must use its chat template.

What a strong answer shows: They know template mismatches silently degrade fine-tuned models.

Intermediate

How do you get reliable structured output from an LLM?

Use the provider's structured output or JSON schema feature, or constrained decoding when self-hosting, so output must match the schema. Then validate in code anyway and handle refusals or truncation, often with one repair retry.

from typing import Literal
from pydantic import BaseModel, ValidationError

class Triage(BaseModel):
category: Literal["billing", "bug", "account", "other"]
priority: Literal["low", "medium", "high"]
summary: str

try:
triage = Triage.model_validate_json(raw_output)
except ValidationError as err:
triage = retry_with_error(raw_output, err)

What a strong answer shows: They rely on schema enforcement plus validation and know max-token truncation breaks JSON.

How does prompt caching work, and how do you structure prompts to benefit?

Providers and engines reuse the computed KV state for a prompt prefix they have seen recently, cutting cost and time to first token. Caches match exact prefixes, so put stable content first (system instructions, tool definitions, reference documents) and variable content last. Cache lifetimes and pricing differ by provider, so check the documentation.

What a strong answer shows: They know one changed token early in the prompt invalidates everything after it.

When do you fine-tune instead of using retrieval or better prompting?

Prompting and retrieval come first: they are cheaper and keep knowledge updatable. Fine-tune to change behavior, such as a strict output style, a domain-specific classification task, shorter prompts at high volume or getting a small model to match a larger one on a narrow task. Fine-tuning is a poor way to inject frequently changing facts.

What a strong answer shows: They separate knowledge problems (retrieval) from behavior problems (fine-tuning).

Explain LoRA and the settings you would choose first.

LoRA freezes the base weights and trains small low-rank matrices added to selected layers, so you train and store a fraction of the parameters. QLoRA does the same over a 4-bit quantized base to save memory.

from peft import LoraConfig, get_peft_model

config = LoraConfig(
r=16, lora_alpha=32, lora_dropout=0.05,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
task_type="CAUSAL_LM",
)
model = get_peft_model(base_model, config)
model.print_trainable_parameters()

What a strong answer shows: They explain rank and target modules as capacity trade-offs and know adapters can be swapped per tenant at serving time.

What makes a good fine-tuning dataset?

Quality over volume: a few thousand clean, consistent, representative examples beat noisy bulk data. Deduplicate, format with the model's chat template, hold out an eval split with no overlap and include edge cases. Check general capability after training, since narrow tuning can degrade other skills.

What a strong answer shows: They check for contamination and regressions, not just training loss.

You use an LLM as a judge in your evals. How do you know the judge is trustworthy?

Label a sample by hand and measure agreement. Use specific rubrics, randomize order in pairwise comparisons to counter position bias, watch for preference toward longer answers and re-check the judge when you change its model or prompt.

from sklearn.metrics import cohen_kappa_score

kappa = cohen_kappa_score(human_labels, judge_labels)

What a strong answer shows: They validate the judge like any other model, with a human baseline.

How would you measure hallucination in a document summarization feature?

Measure faithfulness against the source: break summaries into atomic claims and check each is supported, using a judge model or entailment classifier validated against human labels. Track the unsupported-claim rate on a fixed set, and require cited spans for high-stakes summaries so claims can be checked automatically.

What a strong answer shows: They define a measurable metric instead of saying "we check for hallucinations".

Inference and serving

What is the KV cache, and how much memory does it need?

During decode, the model stores keys and values for every previous token at every layer so it does not recompute them. Memory per token is 2 × layers × KV heads × head dimension × bytes per value.

# 32 layers, 8 KV heads, head_dim 128, fp16 (2 bytes)
per_token = 2 * 32 * 8 * 128 * 2 # 131,072 bytes = 128 KiB
per_seq_8k = per_token * 8192 # 1 GiB per 8k-token sequence

That is why long contexts and high concurrency run out of GPU memory before compute. Grouped-query attention, KV cache quantization and paging all target this cost.

What a strong answer shows: They can do the arithmetic and connect it to batch size limits.

Why does vLLM get much higher throughput than a naive serving loop?

PagedAttention stores the KV cache in fixed-size blocks, avoiding fragmentation and allowing prefix sharing. Continuous batching adds and removes sequences at every step instead of waiting for a whole batch to finish. Together they keep the GPU busy with many concurrent requests.

from vllm import LLM, SamplingParams

llm = LLM(model=MODEL_PATH, tensor_parallel_size=2,
enable_prefix_caching=True, max_model_len=16384)
outputs = llm.generate(prompts, SamplingParams(temperature=0.2, max_tokens=512))

What a strong answer shows: They explain the mechanisms, not just the benchmark claims.

What are the trade-offs of quantizing a model to 8-bit or 4-bit?

Weight quantization (for example AWQ or GPTQ at 4-bit, or FP8 on supporting GPUs) cuts memory and often raises throughput, letting a larger model fit on fewer GPUs. Quality loss varies by model and task, often showing up in reasoning, code or long-tail languages first. Evaluate on your own tasks before and after.

What a strong answer shows: They measure task quality and do not trust generic benchmarks.

How does speculative decoding speed up generation?

A small draft model, or extra prediction heads, proposes several tokens; the large model verifies them in one forward pass and accepts the matching prefix. Output distribution is preserved, and speed improves when the draft is often right. Gains shrink at high batch sizes where the GPU is already saturated.

What a strong answer shows: They know when it helps (latency-sensitive, low concurrency) and when it does not.

How do you tune a serving deployment for latency versus throughput?

Define SLOs for time to first token and time per output token. Larger batches and more concurrent sequences raise throughput but slow each request. Tune max concurrent sequences, chunked prefill, which keeps long prompts from stalling decoding, and replica count against load tests that use your real prompt and output lengths.

What a strong answer shows: They test with realistic length distributions, not a single prompt.

A model does not fit on one GPU. What are your options?

Quantize it, use tensor parallelism to split each layer across GPUs in one node (needs fast interconnect), or pipeline parallelism to split layers across stages or nodes. Tensor parallelism lowers latency; pipeline parallelism scales further but adds bubbles.

What a strong answer shows: They reason about interconnect bandwidth and latency together.

Estimate GPUs needed to serve 50 requests per second with 2,000 input and 300 output tokens each.

Benchmark one replica on your hardware with that length mix to find sustainable requests per second within SLO, then divide and add headroom for peaks and failures. Sanity-check against KV cache capacity: concurrency equals arrival rate times average request duration, and each in-flight request holds its KV cache.

What a strong answer shows: A method grounded in measurement and Little's law, not a guessed number.

Senior and architecture

How do you choose a model for a new task?

Build a task-specific eval set first, then compare candidate hosted and open-weight models on quality, latency, cost and data constraints. Often a smaller model with good prompting or light fine-tuning matches a larger one on a narrow task. Re-run the comparison as models change.

What a strong answer shows: Model choice is an eval result, not a brand preference.

When does distillation make sense, and what do you check first?

When a large model performs well on a narrow, high-volume task and cost or latency matters, generate training data from it and fine-tune a smaller model. Check the provider's terms on using outputs for training, filter the generated data for quality and compare the student against the teacher on held-out evals.

What a strong answer shows: They check licensing and validate the student independently.

Design an evaluation platform used by ten teams shipping LLM features.

Provide versioned datasets, reusable scorers (exact match, schema checks, validated judges), experiment tracking across prompt, model and parameter changes, CI hooks that block regressions and sampling of production traffic for ongoing scoring. Make results comparable across runs by pinning judge versions.

What a strong answer shows: They make evals cheap to run and hard to skip.

A multi-turn assistant degrades after 40 turns. How do you manage context?

Measure what fills the window: tool results, retrieved text and history. Summarize or compact older turns, drop stale tool output, keep durable facts in structured memory and re-retrieve when needed. Preserve the cache-friendly prefix while trimming the tail.

What a strong answer shows: They treat context as a budget to engineer, not something to fill.

How do you run red-teaming for an LLM deployment?

Maintain an adversarial suite: jailbreaks, prompt injection through retrieved content, attempts to extract system prompts or other users' data and harmful-content probes relevant to the domain. Run it on every model or prompt change, track pass rates and add new failures from production reports.

What a strong answer shows: Red-teaming is continuous and regression-tested.

When should a team self-host open-weight models at all?

When data cannot leave its environment, when steady high volume makes GPUs cheaper than tokens or when deep customization is required. Budget for serving engineers, GPU capacity, upgrades and evals; hosted APIs carry that load for you.

What a strong answer shows: They count the full operating cost.

Red flags to watch for

A practical exercise

Run a three-hour take-home. Provide 300 labeled support tickets and access to a small open-weight model (runnable locally or through a provided endpoint) plus a hosted model. Ask the candidate to build a classifier that returns validated structured output, compare prompting the hosted model with a LoRA fine-tune or few-shot prompt on the small model, and report quality, latency and cost per 1,000 tickets.

Hire senior LLM engineers vetted with these questions

Ryz introduces senior LLM engineers who have fine-tuned models, built eval harnesses and served models in production, working with Anthropic, OpenAI and open-weight models. They are the top 1% of the candidates we interview, they work on your team and repos, and they keep hours within ±1h of US time zones. Learn how we screen in our vetting process, or use our LLM engineer job description as a template.

FAQ

Do LLM engineers need to know how to train models from scratch?

Rarely. Pretraining is done by a handful of labs. Most LLM engineering work is adaptation, evaluation and serving, so test fine-tuning, inference and eval skills instead.

How do I test LLM skills without giving candidates GPUs?

Use small open-weight models that run on a laptop, or a provided endpoint, and focus on method. Memory and capacity estimates can be tested on a whiteboard with given model dimensions.

How fast does LLM interview content go out of date?

Specific model names and prices change monthly, so avoid them in questions. Concepts like KV cache, tokenization, LoRA and eval design stay stable and test durable skill.

Questions we didn't answer? Email info@ryzlabs.com.

Explore Ryz Labs

Staff augmentationDedicated development teamsAI pod teamsForward deployed engineersNearshore software developmentAI engineering teamsHire engineers by roleRyz Labs vs competitorsAlternatives guidesBuyer guidesCase studiesHow we vet engineers
Ryz Labs

Senior engineers in your time zone. AI pod teams that ship.

Tell us who you need. You'll get a scoped plan, a price and the names of the people who would do the work.

Start a conversation →