Hiring for this role?

Start free with this plan

Free for your first open role.

Why Hirezen?
  • Every interviewer runs the same script and marks the same signals.
  • AI drafts the write-ups, and the debrief puts every read side by side.
  • No ATS to set up first, and no bot in the call.

DevOps Engineer interview questionsTechnical Interview round

A 60 min interview plan with a time-boxed script, what each question is for, and the signals to score against. Key skills: A service two weeks into Kubernetes: CPU throttling under a low average, a shutdown that never drains, five-second DNS lookups and gRPC that skips new pods — container runtime, service networking and live troubleshooting.

Opening

Who is interviewing, how the round will run, and a question to settle the candidate in. The standard opening

A p99 the average hides

16 min
What this part is for

Purpose

Runs over one page you show when the round starts and leave on screen; nothing is sent ahead. Write it to this composition and add nothing else. The service is `quotes`, which prices delivery. The storefront calls it over HTTP through the ingress, which balances requests across pod addresses; `checkout`'s six pods call it over gRPC through the ClusterIP Service `quotes`, each opening one channel at startup and giving every call a 2-second deadline. Every quote makes one call to a carrier's API at `rates.carrier.example` over HTTPS. Two weeks ago `quotes` moved from six 4-vCPU virtual machines, where its evening-peak p50 was 16 ms and p99 70 ms, to a Deployment on 16-vCPU nodes: 12 replicas, rolled out three at a time; an autoscaler holding 12 to 40 at 80% of requested CPU; each pod requesting 500m CPU and 512Mi, limited to 1 CPU and 512Mi. The image is Debian-based; its entrypoint is a four-line shell script that exports two secrets read from files and runs `/app/quotes` on its last line. The service starts a worker thread per CPU it can see and logs `worker threads: 16`; on SIGTERM it logs `draining`, stops accepting connections, finishes in-flight requests within 20 seconds and exits. No preStop hook or grace period is set, so the default 30 seconds applies. The pod's `/etc/resolv.conf`: `search quotes.svc.cluster.local svc.cluster.local cluster.local`, `nameserver 10.96.0.10`, `options ndots:5`. Then three problems. Latency: at the evening peak p50 is 17 ms and p99 390 ms; CPU averages 0.42 cores per pod; every pod is throttled in 15–45% of its CPU periods (`container_cpu_cfs_throttled_periods_total` over `container_cpu_cfs_periods_total`); a pull request raising the autoscaler's minimum to 24 is in review. Rollouts: each one, `checkout` logs 10–20 calls failing with UNAVAILABLE, in bursts about 30 seconds after each batch of old pods starts terminating, and the storefront's error rate does not move; old pods take 30 to 31 seconds to disappear, and none has logged `draining`. Carrier calls: 0.3% take 5.0–5.4 seconds, almost none 1.5–5, and `checkout` logs DEADLINE_EXCEEDED on about 0.3% of its calls at all hours; at peak the service makes 1,100 carrier calls a second and opens as many connections; the cluster DNS answers about 9,000 queries a second, three quarters NXDOMAIN, led by `rates.carrier.example.quotes.svc.cluster.local` and the same name under the two shorter search domains. Two things are fine and are there to be left alone: the memory settings — usage peaks at 230Mi, no pod has been killed for memory — and `checkout`'s deadline. If asked, the carrier client is built inside the request handler, and five pods run hotter than the other seven; say no more about those five before the last section, a live problem whose evidence is listed there. Book 70 minutes; the candidate's questions come after the hour.

•

I'm [YOUR_NAME] and I run the platform our services are deployed on at [COMPANY_NAME]. No coding challenge and no quiz today: one page about a service we moved onto Kubernetes two weeks ago, and what has gone wrong since. The last part is live — something will break, and I'll answer for the cluster.

What this line is for

Purpose

Tells a candidate expecting a coding test or a quiz that the hour goes on one running system, and that the last part is live.

•

Read it through first; take five minutes. Some of what's on it is working exactly as it should.

What this line is for

Purpose

Lets the candidate read in silence and warns, without saying which, that some lines are healthy — so leaving the memory settings and the deadline alone is the candidate's own judgement.

•

Start with latency. The p99 went from 70 to 390 milliseconds in the move, while the pods average under half their CPU limit. Explain that, and tell me whether the pull request doubling the replicas fixes it.

What this question is for, and what to listen for

Purpose

Container runtime, read through a resource limit. The limit, the thread count and the throttling figure sit on separate lines; the read is whether the candidate joins them into one mechanism instead of taking 0.42 cores as headroom.

Signals to score

  • Takes the throttled periods, not the 0.42-core average, as what needs explaining
  • Explains the quota: a 1-CPU limit gives the container's threads 100 ms of CPU between them per 100 ms period, then all of them wait for the next
  • Connects `worker threads: 16` to the node's 16 CPUs, and works out that sixteen busy threads spend the quota in about 6 ms
  • Explains why p50 held and p99 did not: a request finishing while the quota lasts is untouched, one caught when it runs out waits for the next period, sometimes more than once
  • Sets the five-second carrier calls aside: at 0.3% they sit beyond the 99th percentile
  • Says doubling the replicas lowers the average without changing the bursts, for twice the requested CPU
  • Matches parallelism to the CPU allowed — worker threads sized from the limit if they don't block on I/O, or a limit that fits the threads
  • Weighs dropping the CPU limit and keeping an honest request: what the request still guarantees on a busy node, and what an unlimited pod can take from its neighbours
  • Keeps the memory limit: memory cannot be throttled, and a container over its limit is OOM-killed
  • Judges the fix by throttled periods and p99 at the next evening peak, not by average CPU

Follow-up questions

  • The average is 0.42 cores against a limit of 1. Where did the rest of the limit go?
  • Why did the p50 barely move?
  • The pull request doubles the replicas. What does it do to the p99, and what does it cost?
  • Take the CPU limit off. What stops this pod taking a neighbour's CPU?
  • Would you take the memory limit off as well?

Thirty seconds to go

13 min
What this part is for

Purpose

Container runtime again, from the other end of a pod's life, and kept at that layer: the signal path, and what each caller's connection does between Kubernetes deciding a pod should stop and the pod disappearing.

•

Every rollout, `checkout` sees a burst of failed calls, and every old pod takes thirty seconds to go. Tell me what happens to a pod between being told to stop and disappearing, why `checkout` notices when the storefront doesn't, and what you would change.

What this question is for, and what to listen for

Purpose

Tests whether the candidate reads "every pod takes 30 seconds" and "no pod logged `draining`" as one fact, and knows the signal path well enough to see what the obvious fix leaves broken.

Signals to score

  • Reads the 30 seconds and the missing `draining` line together: the handler never ran, and every pod was killed at the default grace period
  • Finds the cause in the entrypoint: the shell is PID 1 and alone receives SIGTERM, which a PID 1 without a handler ignores
  • Ties the bursts to the kill: `checkout`'s long-lived connections stay on the old pod, still serving, until SIGKILL cuts them mid-call
  • Explains why the storefront is spared: the ingress picks a pod per request and stops picking a terminating one within a second or two
  • Fixes the signal path with `exec` on the last line — under a minimal init if the service can't reap children — and knows an init in front of the unchanged script is not the fix: by default tini signals only the shell, and an init that signals the whole group still exits with the shell and kills the service mid-drain
  • Predicts the next failure: the service will stop accepting the moment it hears SIGTERM, while its removal from the endpoints is still reaching the nodes and the ingress
  • Orders the shutdown: a few seconds still serving, then stop accepting, tell gRPC clients to reconnect elsewhere, finish in-flight calls, exit
  • Sizes the grace period from that sequence with room to spare, so the kill never fires in a normal shutdown
  • Proves it with a rollout under load: failed calls per caller, terminations in seconds, `draining` in every old pod's log

Follow-up questions

  • Every old pod takes 30 seconds to go. What is it doing for those 30 seconds?
  • Why does `checkout` see errors when the storefront doesn't?
  • A pull request adds `exec` to the entrypoint's last line. Once it merges, what fails next?
  • How long should the grace period be, and what is it made of?
  • How will you know the next rollout was clean?

Just over five seconds

13 min
What this part is for

Purpose

Service networking, read on the call that leaves the cluster: name resolution, connection reuse, and what a fixed delay in a latency distribution means.

•

Now the carrier calls. Why do some of them take just over five seconds, and what would you change, in what order?

What this question is for, and what to listen for

Purpose

A distribution with a gap between 1.5 and 5 seconds is a timer firing, not a slow server. The read is whether the candidate derives the lookup volume from `resolv.conf` and the call rate, and fixes what causes the lookups before scaling what answers them.

Signals to score

  • Reads calls just over five seconds, with almost none between 1.5 and 5, as a lost packet retried on a fixed timer, not a slow carrier
  • Names the resolver's five-second retry as the likely timer, and knows a lost TCP connection attempt would show at about one second
  • Connects `checkout`'s 0.3% of DEADLINE_EXCEEDED to the same calls, and leaves its deadline alone
  • Counts one lookup from `resolv.conf` — three search domains first, A and AAAA each, eight queries — and matches 1,100 lookups a second to the 9,000 the cluster DNS answers
  • Traces the lookups to a new connection per call, and says reuse removes most of them and a TCP and TLS handshake from every call
  • Lowers the cost of a lookup — a smaller `ndots` in the pod's DNS config, or a fully qualified name with its trailing dot tested against the certificate check and Host header
  • Asks where queries are lost before adding DNS replicas: the DNS pods' throttling and errors, the nodes' connection-tracking counters
  • Proposes a DNS cache on every node, answering repeat queries locally and outside connection tracking
  • Times DNS separately in the client, so the next slow lookup is labelled as one
  • Caps how long a reused connection lives, so the service notices when the carrier's address changes

Follow-up questions

  • Why five seconds, and not one or three?
  • How many DNS queries does one carrier call cost today?
  • Which single change removes the most of those queries?
  • Would you add DNS replicas?
  • Once connections are reused, when does the service notice that the carrier has moved?

Eighteen pods, five busy

18 min
What this part is for

Purpose

Live troubleshooting, the round's only read of it. None of this is on the page: give each item only when the candidate asks for something it answers; anything unlisted shows nothing unusual. Errors: `checkout`'s calls to `quotes` end in DEADLINE_EXCEEDED — about four in ten, against the usual 0.3% — and its other calls are normal; the storefront's requests succeed under a longer timeout, but those served by five particular pods take several seconds and the rest about 20 ms. CPU per pod: those five at 1.0 core, their limit, throttled in nearly every period; the other thirteen, all six new pods among them, at about 0.18. Autoscaler: 18 replicas, 81% of requested CPU against an 80% target, unchanged for four minutes. Requests per pod: ingress traffic even, about 39 a second each; `checkout`'s gRPC calls arrive only at the five — about 250 a second at each of four, 500 at the fifth. `checkout`: six pods started three days ago, each holding one connection to the Service's cluster IP, two of them to the same pod; at last night's peak each connection carried about 40 calls a second, and the five have run hottest at every evening peak since. Capacity: one CPU of `quotes` serves about 220 calls a second. Nodes: 35–50% CPU, no memory pressure, evictions or restarts. Carrier and DNS: the page's shape at the new volume — about 13,000 queries a second, since the busy pods make no more carrier calls than their CPU allows. Changes: nothing deployed for 24 hours; `quotes` last rolled out four days ago. Costs: a `quotes` rollout takes six minutes; `checkout` scales or restarts in two; the cluster can resize a running pod's CPU in place; `checkout`'s gRPC library can balance round-robin across a headless Service, a configuration change and a redeploy.

•

Last part, and it's live. It's 11:08 on a Tuesday. A promotion went out at 11:00, and calls to `quotes` jumped to about twice the usual evening peak, driven by `checkout`. The autoscaler has taken it from 12 pods to 18. Checkout's calls to `quotes` are missing their 2-second deadline — about four in ten — and it isn't improving. Tell me what you look at, and I'll tell you what you see.

What this line is for

Purpose

Opens with what an engineer called in would know — a symptom, a number, one recent change — so that everything else has to be asked for.

•

What's wrong, what do you do in the next fifteen minutes, and what do you change next week so that a surge can't do this again?

What this question is for, and what to listen for

Purpose

Reads how the candidate narrows a live fault — who is failing, where, since when — before touching anything, and whether their first move changes where `checkout`'s calls go or only adds pods those calls will never reach.

Signals to score

  • Establishes who is failing before changing anything: `checkout`'s gRPC calls, while the storefront is slow only on some pods
  • Asks for per-pod numbers, not the average, and finds five pods at their limit beside thirteen nearly idle
  • Reads the autoscaler's 81% as pods at 200% and 36% of request averaged together, and says more replicas cannot help
  • Explains the idle new pods: each `checkout` pod holds one long-lived HTTP/2 connection, routed once when it opened, and every call since has used it
  • Asks what changed, and takes "nothing but traffic" as evidence that the imbalance predates the surge
  • Works out that one connection's 250 calls a second needs more than a pod's one CPU, so moving connections cannot fix it
  • Picks a first move that cuts the load per connection or adds CPU behind it — more `checkout` pods, or more CPU for the busy five — and says what it costs
  • Rules out the moves that only look like action: restarting `checkout`, deleting the busy pods, raising the autoscaler's maximum, a longer deadline
  • Confirms the first move on deadline failures and per-pod CPU before making a second
  • Fixes it next week with per-call balancing — round-robin over a headless Service, or a proxy that balances requests — and a maximum connection age on the server

Follow-up questions

  • Twelve pods became eighteen. Why are the six new ones idle?
  • The autoscaler says 81% against a target of 80%. Is it wrong?
  • You restart all six `checkout` pods. Where do their calls go?
  • What is your first move, and how will you know within five minutes whether it worked?
  • What would have shown you this before the promotion?

Closing

That's the hour. What do you want to know about how our services run — how we set limits, what a rollout costs us, what last went wrong in production?

What this line is for

Purpose

Not scored. The close is the candidate's time to learn what they would be running; anything worth keeping goes in the notes.

And so you know what you'd be walking into: [tell the candidate one true, specific weakness in how your services run — CPU limits copied from a template and never measured, a service that has never shut down cleanly, dashboards that show only averages]. Fixing it would be part of the job.

What this line is for

Purpose

A true, unflattering detail about your own production tells the engineer this round is for that the job is real, and warns off one who wanted a finished platform. Check that it is still true on the day.

Their questions for you, and what happens next. The standard closing

DevOps Engineer interviews — common questions

Who is this DevOps Engineer interview plan for?
It is written for the interviewer, not the candidate: the hiring manager, engineer or panel member running the Technical Interview round for a DevOps Engineer role. It gives you a 60 min script to follow in the conversation — 4 questions with what each one is for and the signals to score against — so you are not writing the round from scratch the night before.
What does the Technical Interview round assess?
This round is focused on: A service two weeks into Kubernetes: CPU throttling under a low average, a shutdown that never drains, five-second DNS lookups and gRPC that skips new pods — container runtime, service networking and live troubleshooting. It works through A p99 the average hides, Thirty seconds to go, Just over five seconds and Eighteen pods, five busy, scoring against 39 observable signals, with follow-up prompts on all 4 questions for going deeper where an answer is thin.
How is the 60 min split up?
60 min on 4 questions. The questions take in A p99 the average hides (16 min), Thirty seconds to go (13 min), Just over five seconds (13 min) and Eighteen pods, five busy (18 min). The timings are there so the round stays on schedule and every candidate gets the same shape of interview — which is what makes two candidates comparable afterwards.

Hiring for this role?

Open this plan in Hirezen and make it a position in one click.

  • Every interviewer runs the same script and marks the same signals.
  • AI drafts the write-ups, and the debrief puts every read side by side.
  • No ATS to set up first, and no bot in the call.
Start free with this plan

Free for your first open role.