Machine learning engineer job description template (2026)
A complete machine learning engineer job description you can copy, plus seniority levels and tips for hiring someone who gets models into production and keeps them there.
By the Ryz Labs team · Updated October 2026
A machine learning engineer job description should say which models the person will put into production and what happens to them after launch. Ranking, recommendations, fraud scoring, forecasting and computer vision each have different data, latency and monitoring needs. Spell out whether the role trains models, builds the platform other people train on, or both, and name your training and serving stack. The template below is written for a senior ML engineer who owns the path from training data to a monitored production model, including feature pipelines, serving and retraining. If the work is mostly LLM features in a product, an AI engineer or LLM engineer description may fit better.
Machine learning engineer job description template
Job title
Senior Machine Learning Engineer (Training, Serving and MLOps)
Employment type: full-time or contract. Location: remote, with at least four hours of overlap with US Eastern time.
About the role
We are looking for a senior machine learning engineer to build and run the models behind [ranking / recommendations / fraud detection / forecasting] in [product name]. Our models are trained on [Databricks / SageMaker / Vertex AI / Kubernetes] and served [online through an API / in batch jobs], with features from [Feast / Tecton / an in-house store]. You will own models from data to production, work closely with data scientists and backend engineers, and report to [title].
Responsibilities
- Build training pipelines that are reproducible from code, data version and config, with experiment tracking in MLflow or Weights and Biases.
- Design features and build feature pipelines with point-in-time correct joins, so training data matches what the model sees online.
- Run and maintain a feature store or equivalent, keeping offline and online feature values consistent.
- Train and tune models with PyTorch, XGBoost, LightGBM or scikit-learn, choosing the simplest model that meets the target.
- Serve models online with low latency using FastAPI, Triton, BentoML, KServe or a managed endpoint, and in batch where that is cheaper.
- Set up shadow deployments, canaries and A/B tests so new models are proven against the current one before full rollout.
- Monitor production models for data drift, prediction drift, training-serving skew and business metric changes, with alerts that reach an owner.
- Automate retraining and model registry workflows, including approval steps and rollback to a previous version.
- Optimize inference cost and latency with batching, quantization, caching or ONNX export where it helps.
- Work with data scientists to turn notebook prototypes into tested, production-grade code.
- Document models with model cards covering data, metrics, known limits and owners.
Requirements
- 5+ years in software or ML engineering, with at least 3 years shipping and operating ML models in production.
- Strong Python and software engineering habits: tests, typing, code review and packaging.
- Hands-on experience with PyTorch or gradient boosting libraries, and a solid grasp of evaluation metrics, overfitting and leakage.
- Experience building feature pipelines and preventing training-serving skew.
- Production experience serving models online or in batch, with latency and throughput targets.
- Experience with an ML platform such as SageMaker, Vertex AI, Databricks or Kubeflow, plus experiment tracking and a model registry.
- Comfortable with Docker, CI/CD and at least one major cloud provider.
- Experience monitoring models after launch and diagnosing why performance changed.
- Clear written English for design docs, model cards and incident notes.
Nice to have
- Feature stores such as Feast or Tecton.
- Distributed training with Ray, PyTorch DDP or FSDP.
- Streaming features with Kafka and Flink or Spark Structured Streaming.
- Recommendation or ranking systems with two-tower retrieval and learning-to-rank.
- Model monitoring tools such as Evidently, Arize or WhyLabs.
- GPU inference optimization with TensorRT or ONNX Runtime.
Tech stack
Python 3.12, PyTorch, LightGBM, Feast, Databricks, MLflow, Airflow, Kafka, FastAPI, Triton Inference Server, Kubernetes, AWS, Evidently, Grafana. Replace this with your real stack, including your model registry and serving layer.
What success looks like in 6 months
- You have shipped at least one new or retrained model to production through a canary or A/B test, with a measured result on our own metrics.
- Every production model has an owner, monitoring for drift and skew, and a documented rollback path.
- Training runs for your models are reproducible from a commit and a data version, not a notebook on someone's laptop.
- Data scientists can hand off a model and see it in production in days rather than months.
How to apply and interview process
Send your resume or LinkedIn profile and a short note about a model you put into production and what happened after launch. Our process has four steps: a 30-minute intro call, a technical conversation about ML systems you have built, a practical ML system design or coding exercise, and a final conversation with the team you would join. We aim to give feedback within a few days of each step.
Junior vs mid vs senior machine learning engineer
Training a model is the easy part. Seniority in ML engineering shows in how well models behave months after launch.
| Level | Scope | Typical experience | Key skills |
|---|
| Junior | Experiments, feature work and pipeline fixes within an existing ML system | 0-2 years | Python, scikit-learn or PyTorch basics, evaluation metrics, SQL, experiment tracking |
| Mid-level | Owns one model from features to production serving and monitoring | 2-5 years | Feature pipelines, model serving, A/B testing models, drift monitoring, Docker and CI |
| Senior | ML system design, platform choices, standards for training, release and monitoring | 5+ years | Training-serving consistency, feature stores, distributed training, inference optimization, MLOps design |
Tips for writing a machine learning engineer job description that attracts senior talent
- Name the model types and the decision they drive. "Real-time fraud scoring under 50 ms" or "daily demand forecasts for 2,000 stores" tells candidates the real problem. Use your own targets.
- Say where the line with data science sits. If data scientists build models and ML engineers productionize them, say so. If one person does both, say that too.
- Describe your MLOps maturity. Some teams need someone to build a platform from scratch; others need someone to use a mature one well. Both are real jobs.
- List the serving pattern. Online, batch and streaming inference need different skills. Senior engineers filter on this.
- Be honest about GPUs and budget. If training is constrained by compute, mention it. Candidates who enjoy efficiency work will be drawn in.
- Do not turn it into an LLM role by accident. Adding "LLMs, agents and RAG" to a classic ML job attracts the wrong candidates. Mention them only if they are part of the work.
- Ask about post-launch work in interviews. Questions about drift, skew and rollback reveal more than questions about model architecture.
Skip the job post: hire a vetted senior machine learning engineer
Senior ML engineers who have run models in production are scarce, and hiring cycles stretch for months. Ryz Labs can match you with senior machine learning engineers from Latin America who work on your team, work in your repos, ML platform and standups, and keep hours within ±1h of US time zones. Only the top 1% of the engineers we interview make it through our vetting, which covers ML fundamentals, production serving, feature pipelines and MLOps.
Our staff augmentation model lets you add one ML engineer or several. Ryz engineers work on your team, reporting to your leads. Talk to us to scope your team. If you need a whole AI system built, Ryz AI pod teams, dedicated pods of senior engineers that include a tech lead, ML and backend engineers, build it inside your cloud and repos alongside your team. Hiring on your own? Our machine learning interview questions cover what we test.
FAQ
What is the difference between an ML engineer and an AI engineer?
ML engineers usually train, serve and monitor their own models, such as ranking, fraud or forecasting models, and own the MLOps around them. AI engineers mostly build product features on top of foundation models through APIs, focusing on retrieval, evaluation and integration. Pick the title that matches where the work happens.
What should a machine learning engineer job description include?
The model types and use cases, training and serving stack, data sources, latency or throughput targets, monitoring expectations, how ML engineers work with data scientists and what success looks like after six months.
Should I require a PhD for a machine learning engineer?
Rarely. Production ML engineering depends more on software engineering, data handling and operations skills than on research. Require a PhD only if the role includes novel research, and even then consider equivalent experience.
Questions we didn't answer? Email info@ryzlabs.com.