Interview questions

DevOps interview questions for senior engineers (2026)

Practical DevOps questions on pipelines, infrastructure as code, observability and on-call, with notes on what a senior answer sounds like.

These 28 DevOps interview questions focus on how a senior engineer gets code from a merged pull request to production safely and keeps it healthy there: CI/CD design, infrastructure as code with Terraform or OpenTofu, release strategies, observability, incident response and delivery metrics. They are cloud-agnostic on purpose. At Ryz, DevOps candidates are sourced by recruiters, assessed in structured NTRVSTA AI interviews built around scenarios like these, and reviewed by recruiters at each step; the AI scores inform the decision, and people make it.

How to use these questions

DevOps titles cover very different jobs, so decide first whether you need a pipeline and platform builder, an infrastructure-as-code specialist or someone who will own reliability and on-call. Pick from the matching sections.

Ask for the last time something went wrong. Engineers who have run production describe specific failures, the signal that exposed them and what they changed afterward.

Fundamentals

Describe the stages of a CI/CD pipeline you would set up for a typical web service.

On every pull request: lint, type-check, unit tests, build an artifact once, scan dependencies and the image, run integration tests. On merge: publish the immutable, versioned artifact, deploy it to staging, run smoke tests, then promote the same artifact to production with a progressive rollout and automated health checks.

What a strong answer shows: build once, promote the same artifact everywhere, and keep the pull request path fast enough that people don't bypass it.

What is Terraform state, and why does remote state with locking matter?

State maps resources in code to real infrastructure IDs and caches attributes. Without a shared backend, two engineers applying at once can corrupt it or create duplicates. Store it remotely, encrypted, versioned and locked. Recent Terraform releases lock natively in S3, replacing the older DynamoDB lock table.

terraform {
backend "s3" {
bucket = "acme-tfstate-prod"
key = "network/terraform.tfstate"
region = "us-east-1"
encrypt = true
use_lockfile = true
}
}

What a strong answer shows: they know state can contain secrets in plain text and restrict access to it accordingly.

Compare rolling, blue-green and canary deployments, and where feature flags fit.

Rolling replaces instances gradually and is cheap but mixes versions during the rollout. Blue-green runs a full second environment and switches traffic at once, giving instant rollback at double the capacity. Canary sends a small share of traffic to the new version and widens it as metrics stay healthy. Feature flags separate deploying code from releasing behavior, so risky features roll out per user or tenant.

What a strong answer shows: they note that every strategy requires backward-compatible database changes, since two versions run against one schema.

Write a production Dockerfile for a Node.js service and explain the choices.

Multi-stage build so compilers and dev dependencies stay out of the runtime image, a pinned base, dependency install before copying source to keep the layer cache useful, and a non-root user.

FROM node:22-slim AS build
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci
COPY . .
RUN npm run build && npm prune --omit=dev

FROM node:22-slim
WORKDIR /app
ENV NODE_ENV=production
COPY --from=build --chown=node:node /app/node_modules ./node_modules
COPY --from=build --chown=node:node /app/dist ./dist
USER node
CMD ["node", "dist/server.js"]

What a strong answer shows: they mention pinning the base by digest in production and a .dockerignore that keeps .git and local env files out.

How should a pipeline authenticate to the cloud and handle secrets?

Prefer short-lived credentials through OIDC federation from the CI provider, scoped per repository and environment. Application secrets live in a secrets manager and are read at runtime, not baked into images or printed in logs. Protected environments with required reviewers gate production credentials.

What a strong answer shows: they treat CI as a high-value target and limit which branches and workflows can reach production credentials.

What are the DORA metrics, and how do you use them without turning them into targets?

Deployment frequency, lead time for changes, change failure rate and the time to recover from a failed deployment. Together they balance speed against stability. Use them to spot bottlenecks at the team level and track trends, never to rank individuals, because any single metric is easy to game.

What a strong answer shows: they explain how they would measure each one from real data, such as deploy events and incident records.

When does configuration management like Ansible still make sense next to immutable images?

Immutable images and containers are the default for application hosts: build, test, replace. Configuration management still fits long-lived machines you can't easily replace, network devices, bootstrapping image builds, and one-off fleet operations.

What a strong answer shows: they pick based on the lifecycle of the machine, not tool loyalty.

Intermediate

The CI pipeline takes 40 minutes and developers are batching changes to avoid it. How do you speed it up?

Profile where time goes first. Then cache dependencies and build layers, cancel superseded runs, split tests across parallel jobs, run only what a change affects in a monorepo, and move slow end-to-end suites to a later gate.

on:
pull_request:
paths-ignore: ["docs/**"]
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v5
- uses: actions/setup-node@v5
with:
node-version: 22
cache: npm
- run: npm ci
- run: npm test

What a strong answer shows: they measure before and after, and they treat slow CI as a delivery problem, not an annoyance.

After moving a database into a Terraform module, the plan wants to destroy and recreate it. What happened and how do you fix it?

The resource address changed, so Terraform sees one resource removed and another added. Declare the move so state is updated in place, and protect stateful resources from accidental destruction.

moved {
from = aws_db_instance.main
to = module.database.aws_db_instance.this
}

# inside the module
resource "aws_db_instance" "this" {
# ...
lifecycle {
prevent_destroy = true
}
}

What a strong answer shows: they read plans carefully, block applies that destroy stateful resources without review, and know import and removed blocks for related cases.

Someone changed a firewall rule by hand in the cloud console. How do you detect and handle drift?

Run scheduled plans (or plan -refresh-only) and alert when they are not empty. Decide per change: codify it if it was a valid hotfix, or revert it by applying. Long term, restrict console write access in production.

What a strong answer shows: they fix the process that allowed drift, not just the drift.

How would you choose between Terraform and OpenTofu for a new platform in 2026?

OpenTofu forked from Terraform after HashiCorp's 2023 license change and is governed under the Linux Foundation with an open-source license. The two remain broadly compatible for HCL and most providers, but have diverged in features: OpenTofu added client-side state encryption, for example, while Terraform's newer features and HCP integrations are its own. Decide on licensing needs, required features, vendor support and what the team already runs.

What a strong answer shows: a calm, criteria-based answer and awareness that migration later is possible but needs testing.

How do you structure infrastructure code for several environments and many teams?

Small root modules per environment and component (network, data, service) keep plans fast and blast radius small. Shared logic lives in versioned modules with pinned versions per environment, so a module change rolls through dev before prod.

What a strong answer shows: they avoid one giant state file and know the cost of too many tiny ones.

Push-based CD from CI or pull-based GitOps: which do you choose?

Push is simple: the pipeline runs the deploy with credentials to the target. GitOps puts an agent such as Argo CD or Flux inside the environment that pulls the desired state from Git and reconciles continuously, which removes cluster credentials from CI and corrects drift.

What a strong answer shows: they discuss how rollbacks, secrets and promotion between environments work in each model.

How do database migrations fit into a continuous deployment pipeline?

Use expand and contract. Add new columns or tables in a backward-compatible migration that runs before the new code deploys, migrate data in batches, switch reads and writes, and remove old structures in a later release. Run migrations as a separate, single-run pipeline step with a lock, not on every app instance at startup.

What a strong answer shows: they make every migration safe for both the old and the new version of the code.

What does supply chain security mean for a CI/CD pipeline?

Pin third-party actions and base images to commit SHAs or digests, limit workflow token permissions, generate an SBOM, sign images (for example with Sigstore cosign) and verify signatures at deploy time. Produce build provenance in line with SLSA so you can prove what source built which artifact.

What a strong answer shows: they reference a real compromise of a popular CI action, where tag-pinned workflows ran malicious code, as the reason tags are not enough.

Observability and incident response

Define an SLO for a checkout API and design alerting for it.

Pick an SLI users feel, such as the share of checkout requests that succeed within 500 ms, and a target like 99.9% over 30 days. Alert on error budget burn rate with a long and a short window, so fast burns page and slow burns open a ticket.

(
sum(rate(http_requests_total{job="checkout",code=~"5.."}[1h]))
/ sum(rate(http_requests_total{job="checkout"}[1h]))
) > (14.4 * 0.001)
and
(
sum(rate(http_requests_total{job="checkout",code=~"5.."}[5m]))
/ sum(rate(http_requests_total{job="checkout"}[5m]))
) > (14.4 * 0.001)

What a strong answer shows: they alert on symptoms users see, not CPU, and explain what the team does when the budget runs out.

How do logs, metrics and traces work together, and where does OpenTelemetry fit?

Metrics tell you something is wrong cheaply, traces show where in a request path it went wrong, and logs give the detail. OpenTelemetry provides vendor-neutral SDKs and the Collector, so instrumentation stays the same when the backend changes. Correlate them with trace IDs in log lines.

What a strong answer shows: they bring up label cardinality and trace sampling as the main cost controls.

You are paged at 2 a.m.: the error rate on the main API jumped sharply. Walk through the first 20 minutes.

Acknowledge, check the dashboard to confirm user impact and scope, and look at what changed: deploys, config, flags, dependency status. If a recent change lines up, roll it back before diagnosing further. Declare an incident and pull in help early if it is not resolved quickly, with someone keeping a timeline.

What a strong answer shows: mitigation before root cause, clear communication, and no heroics alone at 2 a.m.

What goes into a useful postmortem?

Impact, timeline, contributing factors, what went well, and action items with owners and due dates. It is blameless: it asks why the system allowed the mistake. Follow-up actions are tracked in the normal backlog and reviewed until closed.

What a strong answer shows: they have seen action items die in a document and describe how they kept them alive.

On-call engineers receive dozens of pages a week and most need no action. How do you fix it?

Review every page from the last few weeks: delete alerts nobody acts on, convert cause-based alerts into SLO burn alerts, route non-urgent issues to tickets, and add runbooks to what remains. Track pages per shift as a health metric for the rotation.

What a strong answer shows: they treat alert noise as a reliability risk, because tired responders miss real incidents.

How would you make a deploy roll itself back when it is unhealthy?

Define health gates from SLIs (error rate, latency, saturation) compared against the baseline version, run them automatically during a canary with tools such as Argo Rollouts or Flagger, or a cloud provider's deployment service, and abort and revert when they fail.

What a strong answer shows: they compare canary against baseline, not against a fixed threshold that ignores normal traffic changes.

Senior and architecture

Your company deploys a few times a week and many deploys cause incidents. How do you get to safe daily deploys?

Measure the current lead time and failure causes, then shrink batch size: trunk-based development, feature flags, automated tests at the right layers, progressive rollouts and fast rollback. Deploying more often with smaller changes usually lowers risk, but only once the safety net exists.

What a strong answer shows: a sequenced plan, starting with data, and an understanding that speed and stability improve together.

What does a good internal developer platform look like?

Golden paths: a service template that comes with CI, deployment, observability, secrets and alerting wired in, plus self-service for common requests like a new database or environment. Treat it as a product with users, roadmap and feedback, and keep escape hatches for teams with unusual needs.

What a strong answer shows: they measure adoption and developer time saved, and avoid building a portal nobody asked for.

How would you migrate 200 repositories from Jenkins to GitHub Actions?

Inventory pipelines and group them by pattern, build reusable workflows for the common patterns, migrate a pilot group, then move the rest in waves with both systems running briefly. Handle secrets, self-hosted runners and permissions centrally.

What a strong answer shows: they favor reusable workflows over 200 copies of the same YAML.

How do you enforce guardrails on infrastructure changes without blocking every pull request on a human?

Policy as code checks the plan in CI: no public storage, no open SSH, required tags, approved instance types. Hard rules fail the build; soft rules warn.

package terraform.guardrails

deny contains msg if {
some rc in input.resource_changes
rc.type == "aws_security_group"
some rule in rc.change.after.ingress
rule.from_port <= 22
rule.to_port >= 22
"0.0.0.0/0" in rule.cidr_blocks
msg := sprintf("%s opens SSH to the internet", [rc.address])
}

What a strong answer shows: they test policies, document exceptions, and roll out new rules in warn mode first.

How do you know your backups and disaster recovery plan actually work?

Restore regularly into an isolated environment and measure how long it takes against the RTO. Run game days that fail a dependency or region on purpose, keep runbooks current, and protect backups from the same credentials that could delete production.

What a strong answer shows: "we have backups" is never accepted until a restore has been timed.

Who owns cloud cost in a DevOps model, and how do you keep it under control?

The teams that run services own their cost, supported by tagging, per-team dashboards and budget alerts. The platform team makes the cheap path the easy path: autoscaling defaults, ephemeral preview environments that expire, and right-sized templates.

What a strong answer shows: they make cost visible to the people who can change it.

Monorepo or many repositories: how does the choice change CI/CD?

A monorepo needs affected-only builds, remote caching and clear ownership files, but makes cross-cutting changes atomic. Many repositories keep pipelines simple per service but need versioned shared libraries and reusable workflows to avoid drift between them.

What a strong answer shows: they connect the repo layout to build tooling, ownership and release coordination.

Red flags to watch for

A practical exercise

A 90-minute pairing session works well. Provide a small service repository with a slow, fragile pipeline (no caching, secrets as static keys, artifacts rebuilt for production, a migration run on app startup) and a Terraform directory with one environment hard-coded. Ask the candidate to propose a target pipeline, implement the two highest-value fixes, and sketch how they would add a second environment and an SLO-based alert.

Hire senior DevOps engineers vetted with these questions

Ryz introduces senior DevOps engineers who have built pipelines, platforms and on-call practices in production. They are the top 1% of the candidates we interview, they join your team, repos and incident channels, and they work within ±1h of US time zones. Read how our vetting process works, or adapt our DevOps engineer job description for your opening.

FAQ

Should DevOps interviews focus on one cloud provider?

Only if the role is tied to it. Pipelines, infrastructure as code, observability and incident response transfer across providers, and those are harder to learn than a new console. Check provider depth with one or two scenario questions in your stack.

How do I assess on-call and incident skills in an interview?

Run a short tabletop: describe an alert, reveal dashboards and log lines as the candidate asks for them, and watch how they triage, communicate and decide when to roll back. It shows more than any question about tools.

What is the difference between a DevOps engineer and an SRE?

The roles overlap heavily. DevOps roles usually emphasize delivery: pipelines, infrastructure code and developer platforms. SRE roles emphasize reliability: SLOs, capacity, incident response and reducing toil. Many teams hire one person for both, so write the job description around the work, not the title.

Questions we didn't answer? Email info@ryzlabs.com.

Explore Ryz Labs

Staff augmentationDedicated development teamsAI pod teamsForward deployed engineersNearshore software developmentAI engineering teamsHire engineers by roleRyz Labs vs competitorsAlternatives guidesBuyer guidesCase studiesHow we vet engineers
Ryz Labs

Senior engineers in your time zone. AI pod teams that ship.

Tell us who you need. You'll get a scoped plan, a price and the names of the people who would do the work.

Start a conversation →