Interview questions

AI engineer interview questions for senior hires (2026)

Questions for engineers who build AI features into real software: model APIs, retrieval, tool use, evals, guardrails, cost and what happens after launch.

These 26 AI engineer interview questions are for engineers who build AI features into software people use: integrating model APIs, designing retrieval systems, wiring up tool use and agents, writing evals, controlling cost and latency, and adding guardrails so the feature behaves in front of customers. The focus is product engineering with models, not training them. At Ryz, senior AI engineers are evaluated on the same ground: recruiters source engineers who have shipped AI features to production, candidates complete structured NTRVSTA AI interviews on system design and debugging, and recruiters review every candidate before and after. AI scores are advisory; people decide.

How to use these questions

AI engineering titles cover a wide range, so start by asking what the candidate actually shipped and who used it. Then choose questions that match your roadmap: retrieval-heavy assistants, agentic workflows or extraction pipelines. Mix fundamentals with two or three scenario questions from the RAG and agents section.

Strong candidates talk about failure cases, evals and cost per request as naturally as they talk about prompts. Weak ones describe demos. Include a practical build exercise, because the gap between a demo and a reliable feature is exactly what you are hiring for.

Fundamentals

How is an AI engineer different from an ML engineer?

An ML engineer usually trains and serves models on the company's own data. An AI engineer mostly builds applications on top of foundation models from providers such as Anthropic or OpenAI, or open-weight models: prompt and context design, retrieval, tool integration, evals and the product surface around them. The roles overlap when fine-tuning or self-hosting enters the picture.

What a strong answer shows: A clear view of where their own experience sits and when they would bring in ML specialists.

What belongs in a production call to a model API that a prototype usually skips?

Timeouts, retries with exponential backoff for rate limits and transient errors, a cap on output tokens, logging of inputs, outputs, token counts and latency, and the model version pinned in configuration rather than scattered through code.

import os
import anthropic

client = anthropic.Anthropic(max_retries=3, timeout=30.0)

resp = client.messages.create(
model=os.environ["SUMMARY_MODEL"],
max_tokens=800,
system=SUMMARY_PROMPT,
messages=[{"role": "user", "content": ticket_text}],
)
log_usage(resp.usage.input_tokens, resp.usage.output_tokens)

What a strong answer shows: They treat the model as an unreliable remote dependency and instrument it like one.

Why stream model responses to the UI, and what gets harder when you do?

Streaming cuts perceived latency because users see text as soon as the first tokens arrive. It complicates error handling mid-stream, output validation (you cannot check a full answer before showing it), cancellation when the user navigates away and moderation of partial content. Server-sent events are the usual transport.

What a strong answer shows: They know when not to stream, for example when output must be validated before display.

Explain embeddings and vector search to a backend engineer who has not used them.

An embedding model turns text into a vector so that similar meanings sit close together. You embed documents once, store vectors in an index (pgvector, OpenSearch or a dedicated vector database), embed the query at request time and return nearest neighbors by cosine similarity. Approximate indexes such as HNSW trade a little recall for speed.

What a strong answer shows: They mention that query and documents must use the same embedding model and that re-embedding is a migration.

How do you manage prompts in a codebase with several AI features?

Treat prompts as code: version them in the repo, keep them out of string concatenation scattered across handlers, separate instructions from user-supplied content with clear delimiters, and tie every prompt change to an eval run. Log which prompt version produced each response.

What a strong answer shows: Prompt changes go through review and evals like any other change.

What is the Model Context Protocol, and when would you use it?

MCP is an open protocol for exposing tools, data and prompts to model-based applications through a standard server interface. It is useful when several assistants or clients need the same internal tools, so you write the integration once. For a single feature with two tools, direct function definitions are simpler.

What a strong answer shows: They weigh standardization against added moving parts and think about auth on MCP servers.

Intermediate

A support assistant reads customer emails and can issue refunds. How do you defend against prompt injection?

Assume any text from emails, web pages or documents can contain instructions. Limit what the model can do: scoped tools, server-side authorization checks, refund caps and human approval above a threshold. Keep untrusted content clearly marked as data, and never give the model credentials it could leak. Detection classifiers help but are not a boundary.

What a strong answer shows: They rely on permissions and architecture, not on a cleverer system prompt.

An AI feature costs far more per user than planned. How do you bring it down?

Measure cost per request and per feature first. Then route simple requests to a smaller, cheaper model, trim context (fewer retrieved chunks, shorter history), use provider prompt caching for stable prefixes, cache identical responses, move non-interactive work to batch APIs and cap output length.

What a strong answer shows: They attribute cost before cutting and check quality with evals after each change.

A chat answer takes 8 seconds at p95. Where do you look?

Trace each step: retrieval, reranking, any chained model calls and generation. Common fixes are running independent calls in parallel, streaming the final answer, using a smaller model for classification or routing steps, shrinking prompt size and speeding up retrieval. Report time to first token separately from total time.

What a strong answer shows: They trace the whole request, not only the model call.

How do you know an AI feature works before launch?

Build an eval set from real or realistic inputs, including hard and adversarial cases, with expected outcomes. Score with deterministic checks where possible and model-graded rubrics where not, and run it in CI on every prompt or model change.

import pytest

@pytest.mark.parametrize("case", load_cases("evals/refund_intents.jsonl"))
def test_intent_classification(case):
result = classify_intent(case["email"])
assert result.intent == case["expected_intent"]

What a strong answer shows: Evals exist before launch and gate changes, with a pass-rate threshold agreed with the product owner.

Your model provider has an outage during peak hours. What should the product do?

Degrade gracefully. Options include a fallback to a second provider or model behind the same interface, a circuit breaker so requests fail fast, queuing non-urgent work and a clear message in the UI. Fallback models need their own eval results, since prompts rarely transfer perfectly.

What a strong answer shows: They plan for failure and test fallback quality, not just availability.

Where do guardrails go in an AI feature, and what do they check?

On input: PII redaction where required, policy and abuse checks, size limits. On output: schema validation, policy checks, grounding checks against sources for factual features and blocked actions. On tools: authorization and limits enforced in code. Keep guardrails fast and log every block for review.

What a strong answer shows: They layer controls and measure false blocks as well as misses.

What would you log for an AI feature, and what would you avoid logging?

Log a trace per request: prompt version, model, retrieved document IDs, tool calls and results, token counts, latency, cost and user feedback. Avoid storing raw sensitive data longer than needed; redact or hash it and set retention rules with your security team.

What a strong answer shows: They balance debuggability with privacy.

RAG and agents

Design a RAG assistant over 50,000 internal documents.

Ingest and parse documents, chunk them along structure, attach metadata (source, date, permissions) and embed. At query time, rewrite the query if needed, run hybrid search (keyword plus vector), rerank the top results, then generate with citations. Add incremental re-indexing, freshness tracking and evals for both retrieval and answers.

What a strong answer shows: They cover ingestion, permissions and evaluation, not only the query path.

The answer is wrong even though the right document exists. How do you debug it?

Split the problem. Check whether the right chunk was retrieved in the top k; if not, it is a retrieval issue (chunking, embeddings, query wording, filters). If it was retrieved, it is a generation issue (context ordering, conflicting chunks, prompt). Measure retrieval recall on a labeled set so you fix the right half.

What a strong answer shows: They evaluate retrieval and generation separately.

How do you combine keyword and vector search results?

Keyword search catches exact terms such as product codes and names that embeddings blur; vector search catches paraphrases. Reciprocal rank fusion merges the lists without needing comparable scores.

def rrf(rankings, k=60):
scores = {}
for ranking in rankings:
for rank, doc_id in enumerate(ranking, start=1):
scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank)
return sorted(scores, key=scores.get, reverse=True)

What a strong answer shows: They know when hybrid search helps and follow it with a reranker.

Users must only see answers from documents they are allowed to read. How do you enforce that?

Filter at retrieval time using permission metadata synced from the source system, so unauthorized chunks never reach the model. Telling the model to ignore restricted content is not enforcement. Handle permission changes quickly and audit with test users at different access levels.

What a strong answer shows: Access control lives in the retrieval layer, not the prompt.

Walk through how tool use works in a model API.

You describe tools with names, descriptions and JSON schemas. The model returns a structured tool call instead of text; your code validates it, executes the function, and sends the result back so the model can continue. The application, not the model, runs everything.

tools = [{
"name": "get_order_status",
"description": "Look up the shipping status of one order by its ID.",
"input_schema": {
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"],
},
}]
# if resp.stop_reason == "tool_use": run the tool, return a tool_result block

What a strong answer shows: They validate tool inputs and write descriptions the model can act on.

When should you build an agent instead of a fixed workflow?

Use a fixed workflow (prompt chain, router, extraction pipeline) when the steps are known; it is cheaper, faster and easier to test. Use an agent loop when the path depends on what the model discovers, such as research or multi-step troubleshooting, and cap steps, cost and time.

What a strong answer shows: They default to the simplest design that works.

An agent with write access deleted the wrong records in staging. What safeguards should production have?

Least-privilege credentials per tool, dry-run or preview modes, human approval for destructive actions, idempotent operations, limits on batch sizes and a full audit trail of tool calls. Test agents against scripted scenarios before giving them write access.

What a strong answer shows: They design for the agent being wrong sometimes.

Senior and architecture

When would you self-host an open-weight model instead of using a hosted API?

Self-hosting fits strict data residency, very high steady volume where GPU cost beats per-token pricing, or a need to fine-tune deeply. Hosted APIs win on model quality at the frontier, time to ship and operational simplicity. Many teams use both, routing tasks by sensitivity and difficulty.

What a strong answer shows: They count the operational cost of GPUs and serving, not only the token bill.

The business wants invoice data extraction at 99% accuracy. How do you deliver it?

Measure field-level accuracy on real invoices first. Validate outputs against schemas and business rules (totals add up, dates parse), route low-confidence or failed checks to a human review queue and feed corrections back into evals. The system reaches the target; the model alone may not.

What a strong answer shows: They combine model output with validation and human review.

Your provider releases a new model version. How do you migrate?

Run your eval suite on the new version, compare quality, cost and latency, adjust prompts, then shadow or A/B it on live traffic before switching. Keep model versions pinned in config so upgrades are deliberate and rollback is one change.

What a strong answer shows: Model upgrades are a release process, not a string edit.

Legal is worried about sending customer data to a model provider. How do you address it?

Review the provider's data retention and training terms, available zero-retention or regional options and certifications. Minimize what you send, redact identifiers where the task allows and document data flows for the security review. For the most sensitive cases, consider models hosted inside your own cloud.

What a strong answer shows: They engage with compliance early and design data minimization in.

How do you measure whether an AI feature is worth keeping?

Define the business metric before launch, such as resolution rate, time saved per task or tickets deflected without reopening. Run an A/B test or staged rollout, track user corrections and feedback and compare value against cost per request.

What a strong answer shows: They judge features by outcomes, not usage alone.

How would you plan the first AI feature for a team that has never shipped one?

Pick a narrow, high-volume task with a clear success measure and a human in the loop, such as drafting support replies for agents to edit. Build evals early, ship to internal users first and set up tracing and cost tracking before scaling. Expand scope once quality is measured.

What a strong answer shows: They de-risk with scope and measurement.

Red flags to watch for

A practical exercise

Give a three-to-four-hour take-home: build a small question-answering service over a provided set of 200 policy documents, with document-level permissions for two user roles. Supply 30 labeled questions. The candidate returns the service, a retrieval and answer eval script with results, and a README covering cost per query, latency and known failure cases.

Hire senior AI engineers vetted with these questions

Ryz introduces senior AI engineers who have built retrieval systems, agents and evals into production software, and AI pod teams that ship complete AI systems inside your cloud. Our engineers are the top 1% of the candidates we interview, they work on your team and repos, and they keep hours within ±1h of US time zones. See our vetting process, or start from our AI engineer job description.

FAQ

What is the difference between an AI engineer and an LLM engineer?

The titles overlap. In general an AI engineer owns the feature end to end, from API integration to the user experience, while an LLM engineer goes deeper on model behavior: fine-tuning, inference optimization, context handling and evaluation methodology.

Do AI engineers need a machine learning background?

Not a research one. They need solid software engineering plus working knowledge of embeddings, evaluation, sampling and model limitations. Statistics helps when designing evals and reading A/B results.

How do I test AI engineering skills in a live interview?

Pair on a broken RAG or tool-use feature with a small eval set. Watch whether the candidate measures before changing prompts, isolates retrieval from generation and reasons about cost.

Questions we didn't answer? Email info@ryzlabs.com.

Explore Ryz Labs

Staff augmentationDedicated development teamsAI pod teamsForward deployed engineersNearshore software developmentAI engineering teamsHire engineers by roleRyz Labs vs competitorsAlternatives guidesBuyer guidesCase studiesHow we vet engineers
Ryz Labs

Senior engineers in your time zone. AI pod teams that ship.

Tell us who you need. You'll get a scoped plan, a price and the names of the people who would do the work.

Start a conversation →