DevOps Engineer interview questions and practice.
Automates how software is built, tested, deployed and run, owning CI/CD pipelines, infrastructure as code and environments.
No card for the taster. Full interviews are paid one at a time. Nothing renews.
Last reviewed
This page is still being written: no authored question bank for this competency family. The role is fully supported in the interview itself; only the published question bank is outstanding.
What interviewers for DevOps Engineer actually ask
The question bank for this role is still being written. These are the first three competencies in the model the interview is scored against.
Restores service quickly under pressure using a structured approach, communicates status clearly, and drives root-cause analysis to actions that actually get done.
Defines and measures reliability with SLIs and SLOs, uses error budgets to negotiate with product, and designs for graceful degradation.
Manages cloud infrastructure declaratively and repeatably, with review, testing and safe rollout of infrastructure changes.
What they are really assessing
Interviewers rarely score whether you seemed nice. They score against a model like this one, usually without telling you it exists. Each competency has a weak, adequate and strong shape, and the difference is almost always the level of specific detail you volunteer without being asked.
Incident response & blameless postmortems
Restores service quickly under pressure using a structured approach, communicates status clearly, and drives root-cause analysis to actions that actually get done.
- Weak
- Incident stories are 'we restarted it and it came back'; no timeline, no role clarity, no root cause beyond the symptom; postmortems either do not happen or produce no completed actions.
- Adequate
- Describes an incident with a rough timeline, a mitigation and a root cause, and a postmortem with actions, but cannot say how they decided mitigation order or which actions were completed.
- Strong
- Gives a specific incident with timeline, how they chose to mitigate before diagnosing, how status was communicated and to whom, the contributing factors found, and the completed actions with measured effect on MTTR or recurrence.
Reliability engineering: SLOs & error budgets
Defines and measures reliability with SLIs and SLOs, uses error budgets to negotiate with product, and designs for graceful degradation.
- Weak
- Reliability is 'we aim for 99.9%' with no SLI definition or measurement; alerts are on CPU and disk; cannot describe a decision changed by an error budget.
- Adequate
- Has defined SLIs on latency and availability for a service and alerts on them, but SLOs are not tied to user journeys and error budget policy is not enforced.
- Strong
- Explains how they chose user-facing SLIs, set SLO targets with product, moved to burn-rate alerting, and a case where an exhausted error budget paused feature work or changed a design.
Infrastructure as code & automation
Manages cloud infrastructure declaratively and repeatably, with review, testing and safe rollout of infrastructure changes.
- Weak
- Infrastructure is changed in the console; IaC is used only for initial creation; cannot describe drift, state management or a bad infrastructure change and its recovery.
- Adequate
- Manages most infrastructure in Terraform/CloudFormation/Pulumi with code review and modules, but testing is limited to plan output and drift is discovered by accident.
- Strong
- Describes module design, state isolation, policy checks and a staged rollout of infrastructure changes; gives an incident caused by IaC and the guardrail (policy as code, drift detection) added afterwards.
CI/CD & release engineering
Builds delivery pipelines that make deploying safe, fast and boring: automated testing, progressive delivery, rollbacks and supply-chain hygiene.
- Weak
- Deployments are manual or a single script; rollbacks are 'redeploy the old version and hope'; cannot give pipeline duration, deploy frequency or change failure rate.
- Adequate
- Has built pipelines with automated tests and deploys, uses blue/green or canary for key services, but pipelines are slow or flaky and rollback has not been rehearsed.
- Strong
- Gives DORA-style numbers before and after their work, explains canary analysis and automated rollback, how they cut pipeline time, and how they secured artifacts and dependencies.
Observability & monitoring
Instruments systems so problems are detected before users complain and can be diagnosed from telemetry, while keeping alerts actionable and noise low.
- Weak
- Monitoring is host metrics and a dashboard nobody watches; alerts page for non-issues; debugging relies on SSH and grep.
- Adequate
- Has structured logs, metrics and some tracing in place and reduced noisy alerts, but cannot explain how they decided what to alert on or measure alert quality.
- Strong
- Describes an observability redesign: alerting on symptoms with SLO burn rates, tracing that cut a diagnosis from hours to minutes, alert noise reduced by a stated percentage, and on-call load measured and improved.
Cloud architecture, security & cost
Designs cloud environments that are secure by default, scale appropriately and cost what they should, and can show the trade-offs made.
- Weak
- Uses whatever the tutorial suggested; cannot explain IAM least privilege, network segmentation or why the bill is what it is; cost work is 'we turned off some instances'.
- Adequate
- Applies least-privilege IAM, private networking and autoscaling, and has reduced cost with rightsizing or reserved capacity, but cannot quantify savings or explain the security model end to end.
- Strong
- Gives a real architecture with the security boundaries explained, a measured cost reduction with the levers used (rightsizing, storage tiers, spot, architecture change), and a trade-off where they chose cost over performance or vice versa.
Developer enablement & platform thinking
Treats internal developers as customers: builds self-service platforms and paved roads, measures adoption and friction, and says no to bespoke requests wisely.
- Weak
- Sees the role as gatekeeping tickets; cannot describe a tool or process that reduced developer waiting time, or how they gathered developer feedback.
- Adequate
- Has built self-service tooling (templates, pipelines, environments) that teams adopted, but adoption and time saved are not measured and the feedback loop is informal.
- Strong
- Describes a platform capability with adoption numbers and measured friction reduction (lead time, tickets), how they prioritised requests, and a case where they declined a bespoke request and what they offered instead.
Composure & communication under pressure
Stays methodical during outages and high-stakes changes, gives clear and honest status updates, and manages their own and the team's on-call load sustainably.
- Weak
- Describes incidents as chaotic heroics; status updates were absent or reassuring without evidence; on-call burnout is accepted as part of the job.
- Adequate
- Follows an incident process and gives regular updates, but struggles to say how they kept others calm or what they changed to make on-call sustainable.
- Strong
- Gives specific examples of running a calm incident bridge, an update that told leadership bad news plainly, and structural changes (rotation, alert cuts, runbooks) that reduced pages by a stated amount.
Reading the questions is the easy half. Try answering three of them out loud, to someone who follows up.
Try 5 minutes freeWhat your 30 minutes covers
The same shape as a real first-round interview, pitched at mid-level DevOps Engineer and scored throughout.
Warm-up, then Motivation & fit
Build rapport, settle nerves, and get a short walk-through of your background. Why this role, why this employer, and what you are actually looking for.
Your experience
Two or three real situations from your CV in depth: context, what you did, what happened, what you would change.
Pitched at mid-level scope: owns the reliability and delivery tooling for a set of services; runs incidents and postmortems; accountable for SLOs in their area.
Role-specific questions
The core competencies and domain knowledge for the role, with follow-ups on anything vague.
Drawn from this role's domain: running an incident: roles, mitigation before diagnosis, and communication cadence, defining SLIs and SLOs for a user-facing service and alerting on burn rate and designing a CI/CD pipeline with canary deploys and automated rollback, and the rest of the competency model.
Your questions, then Wrap-up
Your questions for the interviewer, and yes, they are assessed. Next steps and a clean finish.
What changes with seniority
The questions barely change between levels. What changes is the answer they will accept.
| Junior | Mid | Senior | |
|---|---|---|---|
| Scope of ownership | Owns defined pipeline, infrastructure or monitoring tasks with review; participates in on-call with backup. | Owns the reliability and delivery tooling for a set of services; runs incidents and postmortems; accountable for SLOs in their area. | Owns platform architecture and reliability standards across multiple teams; accountable for major incidents, cost and security posture of the platform. |
| Tolerance for ambiguity | Handles tickets with some gaps and escalates unclear or risky changes rather than guessing. | Turns a vague reliability or cost goal into a plan; identifies risk and sequences changes safely. | Defines reliability and platform strategy from business needs; makes build-vs-buy and architecture decisions with incomplete information. |
| People leadership | None formally. | Mentors juniors, reviews infrastructure changes, may lead a small project. | Technical lead for the platform team, drives design reviews, mentors across teams, may run the on-call programme. |
| Who they deal with | Own team, developers requesting support, tech lead. | Development teams, security, product managers, cloud vendor support. | Engineering leadership, security and compliance, finance for cloud spend, vendors. |
What your report would say
Every competency above scored from your own answers, the sentence that cost you quoted back, and your weakest answers rewritten the way a strong DevOps Engineer would have said them.
- Describes module design, state isolation, policy checks and a staged rollout of infrastructure changes; gives an incident caused by IaC and the guardrail (policy as code, drift detection) added afterwards.
The format, not a result. Scores on your report come from what you actually said.
Is the AI interviewer realistic? See a full sample report