LLM engineer job description template (2026)
A complete LLM engineer job description you can copy, plus seniority levels and tips for hiring someone who works on the models themselves, not just the API calls.
By the Ryz Labs team · Updated October 2026
An LLM engineer job description should say how deep into the model stack the person will go. Some teams need someone to fine-tune and serve open-weight models in their own cloud; others need someone to build evaluation and context systems that keep dozens of LLM features consistent at scale. Name the models you run, whether you self-host, your GPU budget, and the quality and latency targets that matter. The template below is written for a senior LLM engineer who owns fine-tuning, inference optimization, evaluation harnesses and context engineering for a company that runs LLMs as core infrastructure. If you mainly need product features built on hosted APIs, an AI engineer description is the better fit.
LLM engineer job description template
Job title
Senior LLM Engineer (Fine-Tuning, Inference and Evaluation)
Employment type: full-time or contract. Location: remote, with at least four hours of overlap with US Eastern time.
About the role
We are looking for a senior LLM engineer to own how large language models are adapted, served and evaluated at [company name]. We use [hosted models from Anthropic or OpenAI] alongside [open-weight models such as Llama, Qwen or Mistral] served on [vLLM / TGI / SGLang] in [AWS / Azure / GCP]. You will decide when to prompt, when to fine-tune and when to distill, make inference fast and affordable, and build the evaluation harness every model change must pass. You will work with AI engineers, ML engineers and platform engineers and report to [title].
Responsibilities
- Fine-tune open-weight models with supervised fine-tuning, LoRA or QLoRA, and preference methods such as DPO, using libraries like Hugging Face TRL, Axolotl or torchtune.
- Build and curate training datasets, including synthetic data generation, deduplication, contamination checks and labeling guidelines.
- Serve models with vLLM, SGLang, TGI or TensorRT-LLM, tuning batching, KV cache settings, tensor parallelism and speculative decoding.
- Apply quantization such as AWQ, GPTQ or FP8 and measure the quality cost of each change, not just the speed gain.
- Build an evaluation harness covering task accuracy, regression suites, safety tests and LLM-as-judge graders calibrated against human labels.
- Run evaluations on every prompt, model and data change, with versioned datasets and reports that show where quality moved.
- Own context engineering at scale: prompt templates, retrieval context budgets, long-context behavior, caching and conversation memory across features.
- Distill capabilities from a large model into a compact, cheaper one when traffic and quality targets justify it.
- Track GPU utilization, cost per thousand tokens, time to first token and throughput, and plan capacity with platform engineers.
- Investigate model failures such as hallucination patterns, refusals, tokenizer issues and degraded quality after updates, and fix them at the right layer.
- Stay current on new open-weight releases and evaluate them against our own tasks before anyone proposes a switch.
Requirements
- 5+ years in ML or software engineering, with at least 2 years working directly with transformer models in production.
- Solid understanding of transformer internals: attention, positional encodings, tokenization, KV cache and sampling parameters.
- Hands-on fine-tuning experience with PyTorch and Hugging Face tooling, including parameter-efficient methods.
- Production experience serving LLMs with an inference engine such as vLLM, SGLang or TGI, and diagnosing latency and memory problems.
- Experience building evaluation pipelines and judging when an eval result is meaningful or noise.
- Working knowledge of GPU hardware limits, memory budgets and multi-GPU parallelism.
- Strong Python and software engineering practices, including tests, reproducible experiments and code review.
- Experience with hosted model APIs and the trade-offs of hosted versus self-hosted models on cost, privacy and quality.
- Clear written English for experiment reports, design docs and model release notes.
Nice to have
- Reinforcement learning from feedback or verifiable rewards, such as PPO or GRPO.
- Custom CUDA or Triton kernels, or contributions to inference engines.
- Distributed training with DeepSpeed, FSDP or Megatron-LM.
- Embedding and reranker model training for retrieval.
- Managed fine-tuning and serving on Amazon Bedrock, SageMaker or Azure AI Foundry.
- Published research, open source model releases or eval benchmarks.
Tech stack
Python 3.12, PyTorch, Hugging Face Transformers, TRL, Axolotl, vLLM, Ray, Anthropic and OpenAI APIs, Llama and Qwen open-weight models, Weights and Biases, Kubernetes with NVIDIA GPUs, AWS. Replace this with your real stack and hardware; LLM engineers will ask about GPUs early.
What success looks like in 6 months
- Every model or prompt change goes through an evaluation harness you built, with results reviewed before release.
- You have shipped at least one fine-tuned or distilled model that meets our quality bar at a lower cost or latency, measured on our own traffic.
- Inference metrics such as time to first token, throughput and cost per request are tracked and have improved.
- The team has a written decision guide for when to use prompting, retrieval, fine-tuning or a different model.
How to apply and interview process
Send your resume or LinkedIn profile and a short note about a model you adapted, served or evaluated, and what you measured. Our process has four steps: a 30-minute intro call, a technical conversation about LLM systems you have worked on, a practical exercise on evaluation or inference, and a final conversation with the team you would join. We aim to give feedback within a few days of each step.
Junior vs mid vs senior LLM engineer
LLM engineering seniority shows in judgment about when to change the model, when to change the context, and how to prove either change helped.
| Level | Scope | Typical experience | Key skills |
|---|
| Junior | Runs fine-tuning and eval jobs designed by others, prepares datasets | 0-2 years | PyTorch basics, Hugging Face, prompting, data cleaning, running eval scripts |
| Mid-level | Owns fine-tuning or serving for one model family, adds eval suites | 2-5 years | LoRA and SFT, vLLM tuning, quantization trade-offs, eval design, GPU debugging |
| Senior | Model strategy, inference architecture, evaluation standards across the company | 5+ years (2+ with LLMs) | Transformer internals, preference tuning, distillation, capacity planning, eval methodology |
Tips for writing an LLM engineer job description that attracts senior talent
- State whether you self-host. Fine-tuning and serving open-weight models is a different job from orchestrating hosted APIs. Senior LLM engineers decide on this point first.
- Describe your GPU situation. Say what hardware you have access to, whether it is reserved or on demand, and who controls the budget. Vague answers here lose strong candidates.
- Name the quality problem. "Our extraction model fails on handwritten forms" or "we need the same quality at a fraction of the latency" gives candidates something concrete to solve.
- Show that evaluation is valued. Teams that ship model changes on gut feel struggle to hire senior LLM engineers. Mention existing eval sets or your plan to build them.
- Separate research from engineering. If the role publishes papers or trains base models, say so. If it adapts existing models for production, say that instead.
- Be specific about data access. Fine-tuning depends on data. Say what labeled data exists, what privacy limits apply and whether synthetic data is allowed.
- Do not list every model and framework. Pick the models and inference engine you run today. A list of every lab and library looks like guesswork.
Skip the job post: hire a vetted senior LLM engineer
Engineers who understand LLM internals and have run models in production are among the hardest people to hire. Ryz Labs can match you with senior LLM engineers from Latin America who work on your team, work in your repos, cloud and standups, and keep hours within ±1h of US time zones. Only the top 1% of the engineers we interview make it through our vetting, which covers transformer fundamentals, fine-tuning, inference and evaluation.
Our staff augmentation model lets you add one LLM engineer or several. Ryz engineers join your team and report to your leads. Talk to us to scope your team. If you need a whole AI system built, Ryz AI pod teams, dedicated pods of senior engineers that include a tech lead, ML and backend engineers, build it inside your cloud and repos alongside your team. Hiring on your own? Our LLM engineer interview questions cover what we test.
FAQ
What does an LLM engineer do?
An LLM engineer adapts, serves and evaluates large language models. Typical work includes fine-tuning open-weight models, optimizing inference engines, building evaluation harnesses and designing how context is assembled for model calls across many features. They work closer to the models than AI engineers who build product features on hosted APIs.
When should I fine-tune instead of improving prompts?
Fine-tune when prompting and retrieval have plateaued on a well-defined task, when you need consistent output formats at high volume, or when a compact model could replace an expensive one. A good LLM engineer will prove the gain with evaluations before and after.
Does an LLM engineer need a research background?
Not necessarily. Many strong LLM engineers come from ML engineering or systems backgrounds. They need a solid grasp of transformer internals and evaluation, plus hands-on production experience. Research experience helps for roles that train base models or invent new methods.
Questions we didn't answer? Email info@ryzlabs.com.