These 26 AI engineer interview questions are for engineers who build AI features into software people use: integrating model APIs, designing retrieval systems, wiring up tool use and agents, writing evals, controlling cost and latency, and adding guardrails so the feature behaves in front of customers. The focus is product engineering with models, not training them. At Ryz, senior AI engineers are evaluated on the same ground: recruiters source engineers who have shipped AI features to production, candidates complete structured NTRVSTA AI interviews on system design and debugging, and recruiters review every candidate before and after. AI scores are advisory; people decide.
AI engineering titles cover a wide range, so start by asking what the candidate actually shipped and who used it. Then choose questions that match your roadmap: retrieval-heavy assistants, agentic workflows or extraction pipelines. Mix fundamentals with two or three scenario questions from the RAG and agents section.
Strong candidates talk about failure cases, evals and cost per request as naturally as they talk about prompts. Weak ones describe demos. Include a practical build exercise, because the gap between a demo and a reliable feature is exactly what you are hiring for.
An ML engineer usually trains and serves models on the company's own data. An AI engineer mostly builds applications on top of foundation models from providers such as Anthropic or OpenAI, or open-weight models: prompt and context design, retrieval, tool integration, evals and the product surface around them. The roles overlap when fine-tuning or self-hosting enters the picture.
What a strong answer shows: A clear view of where their own experience sits and when they would bring in ML specialists.
Timeouts, retries with exponential backoff for rate limits and transient errors, a cap on output tokens, logging of inputs, outputs, token counts and latency, and the model version pinned in configuration rather than scattered through code.
import os
import anthropic
client = anthropic.Anthropic(max_retries=3, timeout=30.0)
resp = client.messages.create(
model=os.environ["SUMMARY_MODEL"],
max_tokens=800,
system=SUMMARY_PROMPT,
messages=[{"role": "user", "content": ticket_text}],
)
log_usage(resp.usage.input_tokens, resp.usage.output_tokens)
What a strong answer shows: They treat the model as an unreliable remote dependency and instrument it like one.
Streaming cuts perceived latency because users see text as soon as the first tokens arrive. It complicates error handling mid-stream, output validation (you cannot check a full answer before showing it), cancellation when the user navigates away and moderation of partial content. Server-sent events are the usual transport.
What a strong answer shows: They know when not to stream, for example when output must be validated before display.
An embedding model turns text into a vector so that similar meanings sit close together. You embed documents once, store vectors in an index (pgvector, OpenSearch or a dedicated vector database), embed the query at request time and return nearest neighbors by cosine similarity. Approximate indexes such as HNSW trade a little recall for speed.
What a strong answer shows: They mention that query and documents must use the same embedding model and that re-embedding is a migration.
Treat prompts as code: version them in the repo, keep them out of string concatenation scattered across handlers, separate instructions from user-supplied content with clear delimiters, and tie every prompt change to an eval run. Log which prompt version produced each response.
What a strong answer shows: Prompt changes go through review and evals like any other change.
MCP is an open protocol for exposing tools, data and prompts to model-based applications through a standard server interface. It is useful when several assistants or clients need the same internal tools, so you write the integration once. For a single feature with two tools, direct function definitions are simpler.
What a strong answer shows: They weigh standardization against added moving parts and think about auth on MCP servers.
Assume any text from emails, web pages or documents can contain instructions. Limit what the model can do: scoped tools, server-side authorization checks, refund caps and human approval above a threshold. Keep untrusted content clearly marked as data, and never give the model credentials it could leak. Detection classifiers help but are not a boundary.
What a strong answer shows: They rely on permissions and architecture, not on a cleverer system prompt.
Measure cost per request and per feature first. Then route simple requests to a smaller, cheaper model, trim context (fewer retrieved chunks, shorter history), use provider prompt caching for stable prefixes, cache identical responses, move non-interactive work to batch APIs and cap output length.
What a strong answer shows: They attribute cost before cutting and check quality with evals after each change.
Trace each step: retrieval, reranking, any chained model calls and generation. Common fixes are running independent calls in parallel, streaming the final answer, using a smaller model for classification or routing steps, shrinking prompt size and speeding up retrieval. Report time to first token separately from total time.
What a strong answer shows: They trace the whole request, not only the model call.
Build an eval set from real or realistic inputs, including hard and adversarial cases, with expected outcomes. Score with deterministic checks where possible and model-graded rubrics where not, and run it in CI on every prompt or model change.
import pytest
@pytest.mark.parametrize("case", load_cases("evals/refund_intents.jsonl"))
def test_intent_classification(case):
result = classify_intent(case["email"])
assert result.intent == case["expected_intent"]
What a strong answer shows: Evals exist before launch and gate changes, with a pass-rate threshold agreed with the product owner.
Degrade gracefully. Options include a fallback to a second provider or model behind the same interface, a circuit breaker so requests fail fast, queuing non-urgent work and a clear message in the UI. Fallback models need their own eval results, since prompts rarely transfer perfectly.
What a strong answer shows: They plan for failure and test fallback quality, not just availability.
On input: PII redaction where required, policy and abuse checks, size limits. On output: schema validation, policy checks, grounding checks against sources for factual features and blocked actions. On tools: authorization and limits enforced in code. Keep guardrails fast and log every block for review.
What a strong answer shows: They layer controls and measure false blocks as well as misses.
Log a trace per request: prompt version, model, retrieved document IDs, tool calls and results, token counts, latency, cost and user feedback. Avoid storing raw sensitive data longer than needed; redact or hash it and set retention rules with your security team.
What a strong answer shows: They balance debuggability with privacy.
Ingest and parse documents, chunk them along structure, attach metadata (source, date, permissions) and embed. At query time, rewrite the query if needed, run hybrid search (keyword plus vector), rerank the top results, then generate with citations. Add incremental re-indexing, freshness tracking and evals for both retrieval and answers.
What a strong answer shows: They cover ingestion, permissions and evaluation, not only the query path.
Split the problem. Check whether the right chunk was retrieved in the top k; if not, it is a retrieval issue (chunking, embeddings, query wording, filters). If it was retrieved, it is a generation issue (context ordering, conflicting chunks, prompt). Measure retrieval recall on a labeled set so you fix the right half.
What a strong answer shows: They evaluate retrieval and generation separately.
Keyword search catches exact terms such as product codes and names that embeddings blur; vector search catches paraphrases. Reciprocal rank fusion merges the lists without needing comparable scores.
def rrf(rankings, k=60):
scores = {}
for ranking in rankings:
for rank, doc_id in enumerate(ranking, start=1):
scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank)
return sorted(scores, key=scores.get, reverse=True)
What a strong answer shows: They know when hybrid search helps and follow it with a reranker.
Filter at retrieval time using permission metadata synced from the source system, so unauthorized chunks never reach the model. Telling the model to ignore restricted content is not enforcement. Handle permission changes quickly and audit with test users at different access levels.
What a strong answer shows: Access control lives in the retrieval layer, not the prompt.
You describe tools with names, descriptions and JSON schemas. The model returns a structured tool call instead of text; your code validates it, executes the function, and sends the result back so the model can continue. The application, not the model, runs everything.
tools = [{
"name": "get_order_status",
"description": "Look up the shipping status of one order by its ID.",
"input_schema": {
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"],
},
}]
# if resp.stop_reason == "tool_use": run the tool, return a tool_result block
What a strong answer shows: They validate tool inputs and write descriptions the model can act on.
Use a fixed workflow (prompt chain, router, extraction pipeline) when the steps are known; it is cheaper, faster and easier to test. Use an agent loop when the path depends on what the model discovers, such as research or multi-step troubleshooting, and cap steps, cost and time.
What a strong answer shows: They default to the simplest design that works.
Least-privilege credentials per tool, dry-run or preview modes, human approval for destructive actions, idempotent operations, limits on batch sizes and a full audit trail of tool calls. Test agents against scripted scenarios before giving them write access.
What a strong answer shows: They design for the agent being wrong sometimes.
Self-hosting fits strict data residency, very high steady volume where GPU cost beats per-token pricing, or a need to fine-tune deeply. Hosted APIs win on model quality at the frontier, time to ship and operational simplicity. Many teams use both, routing tasks by sensitivity and difficulty.
What a strong answer shows: They count the operational cost of GPUs and serving, not only the token bill.
Measure field-level accuracy on real invoices first. Validate outputs against schemas and business rules (totals add up, dates parse), route low-confidence or failed checks to a human review queue and feed corrections back into evals. The system reaches the target; the model alone may not.
What a strong answer shows: They combine model output with validation and human review.
Run your eval suite on the new version, compare quality, cost and latency, adjust prompts, then shadow or A/B it on live traffic before switching. Keep model versions pinned in config so upgrades are deliberate and rollback is one change.
What a strong answer shows: Model upgrades are a release process, not a string edit.
Review the provider's data retention and training terms, available zero-retention or regional options and certifications. Minimize what you send, redact identifiers where the task allows and document data flows for the security review. For the most sensitive cases, consider models hosted inside your own cloud.
What a strong answer shows: They engage with compliance early and design data minimization in.
Define the business metric before launch, such as resolution rate, time saved per task or tickets deflected without reopening. Run an A/B test or staged rollout, track user corrections and feedback and compare value against cost per request.
What a strong answer shows: They judge features by outcomes, not usage alone.
Pick a narrow, high-volume task with a clear success measure and a human in the loop, such as drafting support replies for agents to edit. Build evals early, ship to internal users first and set up tracing and cost tracking before scaling. Expand scope once quality is measured.
What a strong answer shows: They de-risk with scope and measurement.
Give a three-to-four-hour take-home: build a small question-answering service over a provided set of 200 policy documents, with document-level permissions for two user roles. Supply 30 labeled questions. The candidate returns the service, a retrieval and answer eval script with results, and a README covering cost per query, latency and known failure cases.
Ryz introduces senior AI engineers who have built retrieval systems, agents and evals into production software, and AI pod teams that ship complete AI systems inside your cloud. Our engineers are the top 1% of the candidates we interview, they work on your team and repos, and they keep hours within ±1h of US time zones. See our vetting process, or start from our AI engineer job description.
The titles overlap. In general an AI engineer owns the feature end to end, from API integration to the user experience, while an LLM engineer goes deeper on model behavior: fine-tuning, inference optimization, context handling and evaluation methodology.
Not a research one. They need solid software engineering plus working knowledge of embeddings, evaluation, sampling and model limitations. Statistics helps when designing evals and reading A/B results.
Pair on a broken RAG or tool-use feature with a small eval set. Watch whether the candidate measures before changing prompts, isolates retrieval from generation and reasons about cost.
Questions we didn't answer? Email info@ryzlabs.com.