These 28 DevOps interview questions focus on how a senior engineer gets code from a merged pull request to production safely and keeps it healthy there: CI/CD design, infrastructure as code with Terraform or OpenTofu, release strategies, observability, incident response and delivery metrics. They are cloud-agnostic on purpose. At Ryz, DevOps candidates are sourced by recruiters, assessed in structured NTRVSTA AI interviews built around scenarios like these, and reviewed by recruiters at each step; the AI scores inform the decision, and people make it.
DevOps titles cover very different jobs, so decide first whether you need a pipeline and platform builder, an infrastructure-as-code specialist or someone who will own reliability and on-call. Pick from the matching sections.
Ask for the last time something went wrong. Engineers who have run production describe specific failures, the signal that exposed them and what they changed afterward.
On every pull request: lint, type-check, unit tests, build an artifact once, scan dependencies and the image, run integration tests. On merge: publish the immutable, versioned artifact, deploy it to staging, run smoke tests, then promote the same artifact to production with a progressive rollout and automated health checks.
What a strong answer shows: build once, promote the same artifact everywhere, and keep the pull request path fast enough that people don't bypass it.
State maps resources in code to real infrastructure IDs and caches attributes. Without a shared backend, two engineers applying at once can corrupt it or create duplicates. Store it remotely, encrypted, versioned and locked. Recent Terraform releases lock natively in S3, replacing the older DynamoDB lock table.
terraform {
backend "s3" {
bucket = "acme-tfstate-prod"
key = "network/terraform.tfstate"
region = "us-east-1"
encrypt = true
use_lockfile = true
}
}
What a strong answer shows: they know state can contain secrets in plain text and restrict access to it accordingly.
Rolling replaces instances gradually and is cheap but mixes versions during the rollout. Blue-green runs a full second environment and switches traffic at once, giving instant rollback at double the capacity. Canary sends a small share of traffic to the new version and widens it as metrics stay healthy. Feature flags separate deploying code from releasing behavior, so risky features roll out per user or tenant.
What a strong answer shows: they note that every strategy requires backward-compatible database changes, since two versions run against one schema.
Multi-stage build so compilers and dev dependencies stay out of the runtime image, a pinned base, dependency install before copying source to keep the layer cache useful, and a non-root user.
FROM node:22-slim AS build
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci
COPY . .
RUN npm run build && npm prune --omit=dev
FROM node:22-slim
WORKDIR /app
ENV NODE_ENV=production
COPY --from=build --chown=node:node /app/node_modules ./node_modules
COPY --from=build --chown=node:node /app/dist ./dist
USER node
CMD ["node", "dist/server.js"]
What a strong answer shows: they mention pinning the base by digest in production and a .dockerignore that keeps .git and local env files out.
Prefer short-lived credentials through OIDC federation from the CI provider, scoped per repository and environment. Application secrets live in a secrets manager and are read at runtime, not baked into images or printed in logs. Protected environments with required reviewers gate production credentials.
What a strong answer shows: they treat CI as a high-value target and limit which branches and workflows can reach production credentials.
Deployment frequency, lead time for changes, change failure rate and the time to recover from a failed deployment. Together they balance speed against stability. Use them to spot bottlenecks at the team level and track trends, never to rank individuals, because any single metric is easy to game.
What a strong answer shows: they explain how they would measure each one from real data, such as deploy events and incident records.
Immutable images and containers are the default for application hosts: build, test, replace. Configuration management still fits long-lived machines you can't easily replace, network devices, bootstrapping image builds, and one-off fleet operations.
What a strong answer shows: they pick based on the lifecycle of the machine, not tool loyalty.
Profile where time goes first. Then cache dependencies and build layers, cancel superseded runs, split tests across parallel jobs, run only what a change affects in a monorepo, and move slow end-to-end suites to a later gate.
on:
pull_request:
paths-ignore: ["docs/**"]
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v5
- uses: actions/setup-node@v5
with:
node-version: 22
cache: npm
- run: npm ci
- run: npm test
What a strong answer shows: they measure before and after, and they treat slow CI as a delivery problem, not an annoyance.
The resource address changed, so Terraform sees one resource removed and another added. Declare the move so state is updated in place, and protect stateful resources from accidental destruction.
moved {
from = aws_db_instance.main
to = module.database.aws_db_instance.this
}
# inside the module
resource "aws_db_instance" "this" {
# ...
lifecycle {
prevent_destroy = true
}
}
What a strong answer shows: they read plans carefully, block applies that destroy stateful resources without review, and know import and removed blocks for related cases.
Run scheduled plans (or plan -refresh-only) and alert when they are not empty. Decide per change: codify it if it was a valid hotfix, or revert it by applying. Long term, restrict console write access in production.
What a strong answer shows: they fix the process that allowed drift, not just the drift.
OpenTofu forked from Terraform after HashiCorp's 2023 license change and is governed under the Linux Foundation with an open-source license. The two remain broadly compatible for HCL and most providers, but have diverged in features: OpenTofu added client-side state encryption, for example, while Terraform's newer features and HCP integrations are its own. Decide on licensing needs, required features, vendor support and what the team already runs.
What a strong answer shows: a calm, criteria-based answer and awareness that migration later is possible but needs testing.
Small root modules per environment and component (network, data, service) keep plans fast and blast radius small. Shared logic lives in versioned modules with pinned versions per environment, so a module change rolls through dev before prod.
What a strong answer shows: they avoid one giant state file and know the cost of too many tiny ones.
Push is simple: the pipeline runs the deploy with credentials to the target. GitOps puts an agent such as Argo CD or Flux inside the environment that pulls the desired state from Git and reconciles continuously, which removes cluster credentials from CI and corrects drift.
What a strong answer shows: they discuss how rollbacks, secrets and promotion between environments work in each model.
Use expand and contract. Add new columns or tables in a backward-compatible migration that runs before the new code deploys, migrate data in batches, switch reads and writes, and remove old structures in a later release. Run migrations as a separate, single-run pipeline step with a lock, not on every app instance at startup.
What a strong answer shows: they make every migration safe for both the old and the new version of the code.
Pin third-party actions and base images to commit SHAs or digests, limit workflow token permissions, generate an SBOM, sign images (for example with Sigstore cosign) and verify signatures at deploy time. Produce build provenance in line with SLSA so you can prove what source built which artifact.
What a strong answer shows: they reference a real compromise of a popular CI action, where tag-pinned workflows ran malicious code, as the reason tags are not enough.
Pick an SLI users feel, such as the share of checkout requests that succeed within 500 ms, and a target like 99.9% over 30 days. Alert on error budget burn rate with a long and a short window, so fast burns page and slow burns open a ticket.
(
sum(rate(http_requests_total{job="checkout",code=~"5.."}[1h]))
/ sum(rate(http_requests_total{job="checkout"}[1h]))
) > (14.4 * 0.001)
and
(
sum(rate(http_requests_total{job="checkout",code=~"5.."}[5m]))
/ sum(rate(http_requests_total{job="checkout"}[5m]))
) > (14.4 * 0.001)
What a strong answer shows: they alert on symptoms users see, not CPU, and explain what the team does when the budget runs out.
Metrics tell you something is wrong cheaply, traces show where in a request path it went wrong, and logs give the detail. OpenTelemetry provides vendor-neutral SDKs and the Collector, so instrumentation stays the same when the backend changes. Correlate them with trace IDs in log lines.
What a strong answer shows: they bring up label cardinality and trace sampling as the main cost controls.
Acknowledge, check the dashboard to confirm user impact and scope, and look at what changed: deploys, config, flags, dependency status. If a recent change lines up, roll it back before diagnosing further. Declare an incident and pull in help early if it is not resolved quickly, with someone keeping a timeline.
What a strong answer shows: mitigation before root cause, clear communication, and no heroics alone at 2 a.m.
Impact, timeline, contributing factors, what went well, and action items with owners and due dates. It is blameless: it asks why the system allowed the mistake. Follow-up actions are tracked in the normal backlog and reviewed until closed.
What a strong answer shows: they have seen action items die in a document and describe how they kept them alive.
Review every page from the last few weeks: delete alerts nobody acts on, convert cause-based alerts into SLO burn alerts, route non-urgent issues to tickets, and add runbooks to what remains. Track pages per shift as a health metric for the rotation.
What a strong answer shows: they treat alert noise as a reliability risk, because tired responders miss real incidents.
Define health gates from SLIs (error rate, latency, saturation) compared against the baseline version, run them automatically during a canary with tools such as Argo Rollouts or Flagger, or a cloud provider's deployment service, and abort and revert when they fail.
What a strong answer shows: they compare canary against baseline, not against a fixed threshold that ignores normal traffic changes.
Measure the current lead time and failure causes, then shrink batch size: trunk-based development, feature flags, automated tests at the right layers, progressive rollouts and fast rollback. Deploying more often with smaller changes usually lowers risk, but only once the safety net exists.
What a strong answer shows: a sequenced plan, starting with data, and an understanding that speed and stability improve together.
Golden paths: a service template that comes with CI, deployment, observability, secrets and alerting wired in, plus self-service for common requests like a new database or environment. Treat it as a product with users, roadmap and feedback, and keep escape hatches for teams with unusual needs.
What a strong answer shows: they measure adoption and developer time saved, and avoid building a portal nobody asked for.
Inventory pipelines and group them by pattern, build reusable workflows for the common patterns, migrate a pilot group, then move the rest in waves with both systems running briefly. Handle secrets, self-hosted runners and permissions centrally.
What a strong answer shows: they favor reusable workflows over 200 copies of the same YAML.
Policy as code checks the plan in CI: no public storage, no open SSH, required tags, approved instance types. Hard rules fail the build; soft rules warn.
package terraform.guardrails
deny contains msg if {
some rc in input.resource_changes
rc.type == "aws_security_group"
some rule in rc.change.after.ingress
rule.from_port <= 22
rule.to_port >= 22
"0.0.0.0/0" in rule.cidr_blocks
msg := sprintf("%s opens SSH to the internet", [rc.address])
}
What a strong answer shows: they test policies, document exceptions, and roll out new rules in warn mode first.
Restore regularly into an isolated environment and measure how long it takes against the RTO. Run game days that fail a dependency or region on purpose, keep runbooks current, and protect backups from the same credentials that could delete production.
What a strong answer shows: "we have backups" is never accepted until a restore has been timed.
The teams that run services own their cost, supported by tagging, per-team dashboards and budget alerts. The platform team makes the cheap path the easy path: autoscaling defaults, ephemeral preview environments that expire, and right-sized templates.
What a strong answer shows: they make cost visible to the people who can change it.
A monorepo needs affected-only builds, remote caching and clear ownership files, but makes cross-cutting changes atomic. Many repositories keep pipelines simple per service but need versioned shared libraries and reusable workflows to avoid drift between them.
What a strong answer shows: they connect the repo layout to build tooling, ownership and release coordination.
A 90-minute pairing session works well. Provide a small service repository with a slow, fragile pipeline (no caching, secrets as static keys, artifacts rebuilt for production, a migration run on app startup) and a Terraform directory with one environment hard-coded. Ask the candidate to propose a target pipeline, implement the two highest-value fixes, and sketch how they would add a second environment and an SLO-based alert.
Ryz introduces senior DevOps engineers who have built pipelines, platforms and on-call practices in production. They are the top 1% of the candidates we interview, they join your team, repos and incident channels, and they work within ±1h of US time zones. Read how our vetting process works, or adapt our DevOps engineer job description for your opening.
Only if the role is tied to it. Pipelines, infrastructure as code, observability and incident response transfer across providers, and those are harder to learn than a new console. Check provider depth with one or two scenario questions in your stack.
Run a short tabletop: describe an alert, reveal dashboards and log lines as the candidate asks for them, and watch how they triage, communicate and decide when to roll back. It shows more than any question about tools.
The roles overlap heavily. DevOps roles usually emphasize delivery: pipelines, infrastructure code and developer platforms. SRE roles emphasize reliability: SLOs, capacity, incident response and reducing toil. Many teams hire one person for both, so write the job description around the work, not the title.
Questions we didn't answer? Email info@ryzlabs.com.