Data engineer job description template (2026)
A complete data engineer job description you can copy, plus seniority levels and tips for hiring someone who builds pipelines and models people can trust.
By the Ryz Labs team · Updated October 2026
A data engineer job description should say where your data comes from, where it lands, who uses it and what breaks today. Batch ELT into Snowflake for analysts, streaming events through Kafka for product features, and lakehouse work on Databricks for machine learning teams all need different people. Name your warehouse, orchestrator, transformation tool and data volumes, and say whether the role owns data modeling or only the plumbing. The template below is written for a senior data engineer who builds and runs ingestion pipelines, warehouse models and orchestration for analytics and product teams. Copy it and replace the bracketed parts.
Data engineer job description template
Job title
Senior Data Engineer (Pipelines, Warehouse and Data Modeling)
Employment type: full-time or contract. Location: remote, with at least four hours of overlap with US Eastern time.
About the role
We are looking for a senior data engineer to build the data foundation behind [product name]. We move data from [Postgres / Salesforce / Stripe / event streams] into [Snowflake / BigQuery / Databricks], transform it with [dbt], and orchestrate it with [Airflow / Dagster / Prefect]. Analysts, finance, product and data science depend on what you build. You will own pipelines and models end to end, set standards for data quality, and report to [title].
Responsibilities
- Build and maintain ingestion pipelines from databases, SaaS APIs and event streams, using managed connectors such as Fivetran or Airbyte where they fit and custom code where they do not.
- Set up change data capture from production databases with Debezium, Postgres logical replication or a managed CDC service.
- Design warehouse models in dbt: staging, intermediate and mart layers, with dimensional or wide-table patterns chosen for how people query.
- Write incremental models, snapshots and slowly changing dimensions that stay correct when late or duplicate data arrives.
- Orchestrate pipelines in Airflow, Dagster or Prefect, with clear dependencies, retries, backfills and SLAs.
- Add data quality checks with dbt tests, Great Expectations or Soda, plus freshness and volume alerts that reach the right owner.
- Tune warehouse cost and performance: clustering and partitioning, warehouse sizing, query pruning and removing unused models.
- Build streaming pipelines with Kafka, Kinesis or Pub/Sub where the business needs data in minutes, not hours.
- Manage access control, PII masking and data retention in line with our security and privacy requirements.
- Document datasets, owners and definitions in a catalog so analysts can find and trust what they need.
- Work with analysts and data scientists to turn repeated ad hoc queries into maintained models.
Requirements
- 5+ years in data or software engineering, with at least 3 years building production data pipelines.
- Expert SQL, including window functions, query plans and modeling for analytical workloads.
- Production experience with a cloud warehouse or lakehouse: Snowflake, BigQuery, Redshift or Databricks.
- Strong dbt experience, including tests, macros, incremental models and project structure for a growing team.
- Experience with an orchestrator such as Airflow, Dagster or Prefect, including backfills and failure handling.
- Strong Python for pipeline code, API integrations and tooling, with tests.
- Understanding of data modeling approaches such as Kimball dimensional modeling and when to break the rules.
- Experience handling schema changes, late-arriving data and idempotent reruns without corrupting downstream tables.
- Clear written English for data contracts, documentation and incident notes.
Nice to have
- Open table formats such as Apache Iceberg or Delta Lake.
- Spark or Flink for large batch or streaming jobs.
- Reverse ETL with Hightouch or Census.
- Semantic layers such as dbt Semantic Layer, Cube or LookML.
- Data contracts between application teams and the data platform.
- Terraform for warehouse and cloud resources.
Tech stack
Snowflake, dbt Core, Dagster, Fivetran, Debezium, Kafka, Python 3.12, Great Expectations, Terraform, AWS (S3, MSK), Looker, GitHub Actions. Replace this with your actual stack; data engineers decide quickly based on warehouse and orchestrator.
What success looks like in 6 months
- Critical pipelines have freshness and quality alerts with clear owners, and stakeholders hear about data problems from us first.
- You have shipped at least one new source or domain model that a team now uses for weekly decisions.
- Warehouse cost or pipeline runtime has dropped in a way you can show from our own usage data.
- The dbt project has documented conventions, and new models follow them without reminders.
How to apply and interview process
Send your resume or LinkedIn profile and a short note about a pipeline or model you built and kept running. Our process has four steps: a 30-minute intro call, a technical conversation about data systems you have worked on, a practical SQL and modeling exercise on a realistic dataset, and a final conversation with the team you would join. We aim to give feedback within a few days of each step.
Junior vs mid vs senior data engineer
Junior data engineers move data. Senior data engineers decide how data should be shaped, owned and trusted across the company.
| Level | Scope | Typical experience | Key skills |
|---|
| Junior | Adds models and connectors inside an existing project, fixes failed runs with guidance | 0-2 years | SQL, basic Python, dbt models and tests, reading orchestrator logs |
| Mid-level | Owns pipelines for a domain end to end, from ingestion to marts and alerts | 2-5 years | Incremental models, CDC, orchestration, warehouse tuning, data quality checks |
| Senior | Platform architecture, modeling standards, cost governance and data contracts across teams | 5+ years | Dimensional modeling, streaming vs batch trade-offs, lakehouse formats, governance, mentoring |
Tips for writing a data engineer job description that attracts senior talent
- Name the warehouse and orchestrator first. Snowflake with dbt and Dagster and Hadoop with cron jobs are different careers. Senior data engineers screen roles by this pairing.
- Describe your data volumes honestly. Rows per day, number of sources and freshness targets tell candidates what scale they will work at. Use real numbers from your system.
- Say who owns modeling. Some companies split analytics engineering from data engineering. If this person owns dbt marts and metric definitions, say so clearly.
- Mention the mess. Untested models, a single huge DAG or a stalled migration are real parts of the job. Experienced engineers prefer to hear about them up front.
- List the consumers. Finance reporting, product analytics and ML features each create different correctness and latency pressure.
- Do not require every tool. Asking for Spark, Flink, Kafka, Airflow, dbt, Snowflake and Databricks at once reads as if no stack has been chosen. Require what you run.
- Explain data quality ownership. Say whether there is an on-call for data incidents and who answers when a dashboard is wrong.
Skip the job post: hire a vetted senior data engineer
A long hiring cycle for a data engineer means months of stale dashboards and manual exports. Ryz Labs can match you with senior data engineers from Latin America who work on your team, work in your repos, warehouse and standups, and keep hours within ±1h of US time zones. Only the top 1% of the engineers we interview make it through our vetting, which covers SQL, data modeling, orchestration and pipeline reliability.
Our staff augmentation model lets you add one data engineer or several. Ryz engineers work on your team and report to your leads. Talk to us to scope the team you need. If your data work is part of a larger AI system you need built, Ryz AI pod teams, dedicated pods of senior engineers that include a tech lead, ML and backend engineers, build it inside your cloud and repos alongside your team. If you are hiring on your own, our data engineer interview questions cover what we test.
FAQ
What is the difference between a data engineer and an analytics engineer?
Data engineers focus on ingestion, orchestration, infrastructure and reliability. Analytics engineers focus on transforming data into clean, documented models and metrics, usually in dbt. In many teams one person does both, so your job description should say which side carries more weight.
Should a data engineer job description require Spark?
Only if your workloads need it. Many modern data stacks run entirely on a cloud warehouse with dbt and never touch Spark. Requiring it when you do not use it filters out strong warehouse-focused engineers.
What programming languages should a data engineer know?
SQL and Python cover most roles. Scala or Java matter for Spark and Flink heavy work, and Go appears in some streaming platforms. List the languages your codebase uses today rather than every language a data engineer might touch.
Questions we didn't answer? Email info@ryzlabs.com.