Hiring for this role?

Start free with this plan

Free for your first open role.

Why Hirezen?
  • Every interviewer runs the same script and marks the same signals.
  • AI drafts the write-ups, and the debrief puts every read side by side.
  • No ATS to set up first, and no bot in the call.

Site Reliability Engineer (SRE) interview questionsTechnical Interview round

A 60 min interview plan with a time-boxed script, what each question is for, and the signals to score against. Key skills: SLOs and error budgets, alerting, and capacity planning, argued from one booking API's numbers: an SLO per journey, burn-rate pages tested against 139 real ones, a release on a two-thirds-spent budget, and sizing for a noon spike.

Opening

Who is interviewing, how the round will run, and a question to settle the candidate in. The standard opening

One number for three journeys

16 min
What this part is for

Purpose

Runs over one page about one service, shown when the round starts and left in view; nothing is sent ahead. Set it out as short tables holding exactly these facts, which agree with each other. Give the candidate seven minutes with it and book 75, so the reading and their questions sit outside the 60. The service is `reservations-api`, behind a restaurant-booking product: diners search for free tables and book in the app and on the website, and restaurants' host stands run the evening on tablets that poll the same API every 60 seconds. It runs on Kubernetes across three zones, in front of one PostgreSQL primary. Traffic: 1.3 billion requests in the last 30 days, 500 a second on average and about 40 at night; the weekly peak is Friday 18:00–21:00 at 1,200 a second — 1,000 availability searches, 180 floor-view polls, 12 bookings, 8 others — and bookings are about 1% of requests at every hour. The app's results screen shows six restaurants, fetches each one's availability in its own request, and draws when all six have returned. Search latency at Friday peak, from the load balancer's logs: p95 310 ms, p99 1,150 ms. The team's latency panel shows the mean of the twelve pods' own p99s, from Prometheus histograms with buckets at 0.1, 0.25, 0.5, 1, 2.5 and 5 s: 640 ms for search in the same hour. 5xx at the load balancer over 90 days: 0.04% of searches and 0.2% of bookings; the app's own telemetry counts 0.5% of searches at Friday peak as failed — its 8-second timeout or a connection error. The draft SLO, written last quarter and never adopted: "99.99% of requests succeed — not 5xx at the load balancer — per calendar month; p99 latency under 300 ms." By that measure each of the last six months came in between 99.95% and 99.97%. Paging alerts and their pages over the last 28 days, 139 in all: `HighCPU`, a pod above 80% CPU for 5 minutes, 71, none needing action; `P99Latency`, the panel above 1 s for 5 minutes, 38, 30 of them between 01:00 and 05:00 while a nightly job rebuilds availability; `PodRestarts`, three in 15 minutes, 12; `ErrorRate`, 5xx above 1% of all requests for 2 minutes, 9; `DBConnections`, above 80% of the database's limit, 6, all on Friday or Saturday evenings; `CertExpiry`, a certificate within 14 days of expiry, 2, at 03:10; `DiskUsage`, the database disk above 85%, 1, with about 40 days of growth left. Deploys run about twice each working day; eight of them set off `ErrorRate`, failing 1.5–3% of requests with 502s for two to four minutes, and the rest failed less, for under two minutes. The ninth `ErrorRate` page was a database failover on Tuesday the 13th, when every request failed for four minutes. On Friday the 2nd, from 19:05 to 20:35, a release of the deposit flow failed 40% of bookings with 500s, and nothing paged for it: the overall 5xx rate rose from about 0.05% to 0.45%, a restaurant group's manager phoned support at 19:50, and the release was rolled back at 20:30. Capacity: at Friday peak the autoscaler runs 12 pods at about 100 requests a second each, aiming for 70% CPU with 6 to 24 pods. Last month's load test replayed a Friday's mix at one pod: search p99 held under 1 s up to 130 requests a second. Each pod opens a fixed pool of 30 database connections, about 4 of them busy on average at Friday peak; the primary accepts 500 connections, 20 taken by other jobs, and runs at 50% CPU at Friday peak, two-thirds of it on searches that missed the cache. Availability is cached per restaurant and date for 60 seconds, in a cache all pods share, and a restaurant's entries are dropped whenever it takes a booking; 85% of Friday searches hit the cache. Restaurant Week opens next Monday at 12:00, when 900 restaurants release discounted tables for the fortnight at once; the release job drops their cache entries as it writes them. Last year's opening, with 600 restaurants, ran at 3.5 times that year's Friday peak in searches and 8 times in bookings for twenty minutes, and search p99 passed 6 s. Today is Monday the 26th. Two things on the page are fine and are there to be left alone: the autoscaler's 70% target, and the tablets' polling.

•

I'm [YOUR_NAME], and I lead site reliability at [COMPANY_NAME]. Treat this page as a service you have just been handed: nobody has agreed what reliable means for it, the people on call are tired of it, and next Monday is its busiest day of the quarter. Take seven minutes with it; then I'll ask four things it needs from you. I care more about which numbers you argue from than which tools you would use.

What this line is for

Purpose

Frames the page as an inheritance with users, a tired rotation and a date, so the hour goes on what to measure, page on and prepare for, not on definitions.

•

Start with the draft SLO. Rewrite it — what you would measure, where, and against what targets — and tell me what the draft, and then your version, would have said about Friday the 2nd.

What this question is for, and what to listen for

Purpose

Separates knowing the vocabulary from designing an SLO. The draft has ordinary faults — one ratio for three journeys, an availability target never met, and a latency line with no scope or window at 300 ms, which is neither met nor a bucket edge — and the 2nd shows what they cost.

Signals to score

  • Splits the SLO by journey, showing that 40% of bookings failing moved the overall rate only to 0.45%
  • Prices the 2nd: about 2% of a month's budget under one 99.9% SLO, about 40% under a booking SLO of 99.5%
  • Sets each target below what the journey achieves — searches run at 99.96%, so about 99.9%; bookings at 99.8%, so 99.5–99.7% — and drops a 99.99% never met
  • Gives latency a scope, a window and thresholds the histograms can count, where the draft's 300 ms is neither met nor a bucket edge
  • Says the panel's 640 ms, a mean of twelve p99s, is not the p99 of anything, and sums histogram buckets across pods instead
  • Works out that about 6% of six-request screens wait on a request slower than the request p99
  • Says what a not-5xx count at the load balancer misses, and uses the app's 0.5% without adopting it raw
  • Says what a calendar month costs — a bad 31st forgiven on the 1st — and uses a rolling window

Follow-up questions

  • On the 2nd, 40% of bookings failed for ninety minutes. What did the overall number do?
  • The panel says 640 ms and the load balancer 1,150 ms. Which is the p99?
  • A results screen waits for six requests. What does your latency target mean for the screen?
  • The app counts 0.5% of searches failing. Which number goes in the SLO?
  • Why not keep 99.99% as something to aim for?

139 pages and a phone call

16 min
What this part is for

Purpose

Alerting, read from the page history. The incidents are evidence about the alerts, not incidents to run — the role's incident rounds read that — so keep to what should have woken someone, and when.

•

You have 139 pages in 28 days and an incident that paged nobody. Tell me what you keep, demote or delete, what pages instead, and show me with numbers what your version would have done on the 2nd, on the 13th and during the worst of the deploys.

What this question is for, and what to listen for

Purpose

The arithmetic of the round. Anyone can say "page on symptoms"; the separator is turning an SLO into thresholds and windows, replaying them against three real events, and keeping the few cause-based alerts that still earn a page.

Signals to score

  • Turns each SLO into burn-rate alerts pairing a long window with a short one, and says what error ratio each fires at
  • Finds a booking alert would have paged about eleven minutes into the 2nd — at 99.5%, paging at 14.4 times over an hour — not at 19:50
  • Shows one alert on all requests could not: 0.45% against 99.9% burns at 4.5, under every paging threshold
  • Lets a deploy blip stop paging while it still spends budget, and turns the 502s into a ticket to fix pod shutdown
  • Deletes `HighCPU` as a page, since a 70% autoscaling target puts single pods past 80% on most busy evenings
  • Keeps cause-based pages only where the symptom comes too late, like a certificate days from expiry with renewal failing
  • Keeps `DBConnections` as a capacity ticket: the autoscaler pressing on the database's limit
  • Notices a ratio-based alert can still page on 40 requests a second, and asks what the nightly rebuild does to searches
  • Replays the 28 days and counts two or three pages, one of them the 2nd

Follow-up questions

  • What error rate does your fastest booking alert fire at, and how long does 40% take to reach it?
  • What would one alert on all requests have done on the 2nd?
  • A deploy fails 2% of requests for four minutes. Who gets paged, and should they?
  • Which cause-based alert would you still wake someone for?
  • At 03:00 there are about 24 bookings a minute. How many failures does it take to page?

Two-thirds of a budget, four days out

12 min
What this part is for

Purpose

SLOs and error budgets again, from the other end: what the budget is for. Every candidate gets the same booking SLO, whatever they chose earlier, so the arithmetic is the same for all of them.

•

Say the booking SLO is 99.5% over a rolling 30 days. With the 2nd, the 13th and the usual background, about two-thirds of this window's budget is gone. The bookings team has fixed the deposit flow and wants to ship it on Thursday, with deposits for parties of six or more, which restaurants have asked for before Restaurant Week.

What this line is for

Purpose

Gives everyone the same budget, and makes the release something people want for a good reason, so holding it back costs something.

•

Do they ship on Thursday? Tell me what the budget says, what you decide and who decides it, and what you would write down so the next one is not settled in a meeting.

What this question is for, and what to listen for

Purpose

Tests whether the candidate uses a budget as a rule agreed in advance rather than as a veto or a scoreboard. The arithmetic is short; the read is what they do with it four days before the busiest twenty minutes of the quarter.

Signals to score

  • Turns the budget into bookings: about 22,000 left, fewer than the 2nd alone cost
  • Notices the 2nd leaves a 30-day window on Sunday, so Monday's budget will look healthy by the calendar alone
  • Asks what has changed since the 2nd — a test that reproduces it, the booking alert, a fast rollback — before deciding
  • Separates the fix from the new feature, and would ship one without the other
  • Limits the blast radius: a flag per restaurant, a few first, their failure rate watched against everyone else's, and the rate that turns it off agreed beforehand
  • Keeps Monday's capacity work out of any freeze, because it is reliability work
  • Says who decides — product and reliability, agreed in advance, with a named tie-breaker — so the budget sets a default, not a veto
  • Writes the policy down — costly incidents, an exhausted budget, known peaks, exceptions — and has it signed before it is needed

Follow-up questions

  • How many failed bookings are left in this window, and how does that compare with the 2nd?
  • When does the 2nd leave the window, and what will the budget look like then?
  • Restaurants asked for this. Who decides, and what happens if you disagree?
  • Do the capacity changes for Monday stop too?
  • What would you write down so this is not argued again next quarter?

Twelve o'clock on Monday

16 min
What this part is for

Purpose

Capacity planning, from the page's own numbers. The opening is a step, not a curve, and the limits that matter here do not autoscale.

•

Restaurant Week opens next Monday at 12:00. Tell me what load you plan for, what gives out first and at roughly what number, and what you change this week — and what you would not try to change in time.

What this question is for, and what to listen for

Purpose

Most candidates multiply Friday by three and a half and add pods. The separator is finding the database's two limits, connections and CPU; the cache rule that empties itself on the busiest restaurants; and the step at 12:00 that no autoscaler meets.

Signals to score

  • Forecasts a range from last year's multipliers — about 3,500 searches and 96 bookings a second — and what 900 restaurants could add
  • Finds the connection limit: 30 a pod plus 20 reaches 500 at about 16 pods; the autoscaler may go to 24
  • Uses 4 busy connections in 30 to show the pools are oversized, and that busy connections rise if queries slow
  • Finds that at the same hit rate the opening's cache misses alone need more than the whole primary
  • Sees the release job empty the cache for the 900 restaurants at 12:00, and every booking empty it again
  • Sizes staleness to the primary — 12,600 restaurant-dates at a 60-second TTL is about 210 refreshes a second — with jittered expiry, coalesced misses and the booking call deciding
  • Adds read capacity before Monday — a search replica or a larger primary — and prices a resize's failover
  • Pre-scales before 12:00 instead of trusting the autoscaler with a step, and lifts its limits only once the database can take them
  • Wants a load test of the opening's mix, passed or failed on the SLOs rather than on a replayed Friday
  • Has a waiting room, waves of restaurants or a cached-only switch ready for being wrong, and names who may use it

Follow-up questions

  • What do you plan for, and why that number rather than 3.5 times Friday?
  • At how many pods does the database run out of connections?
  • The cache hits 85% on a Friday. Why might it not at 12:00 on Monday?
  • Someone raises the autoscaler's maximum to 48. Does that help?
  • At 12:05 search p99 is four seconds and climbing. What do you have ready?

Closing

That's the page. What do you want to know about reliability here — what our SLOs are and who agreed them, how often the on-call phone rang last week, what we do before our own busiest day?

What this line is for

Purpose

Not scored. A candidate who has carried a pager asks how often it rings and who decides what pages; anything notable goes in the notes.

And so you hear it from me rather than in your first week: [tell the candidate one true, specific weakness in your own reliability work — an alert that pages every night and is ignored, an SLO nobody has agreed, a limit your busiest day will hit]. That would be yours to fix.

What this line is for

Purpose

A specific, true, unflattering fact shows the candidate that this round describes the job, and tells one who wanted it tidy that it is not. Check it is still true before you say it.

Their questions for you, and what happens next. The standard closing

Site Reliability Engineer (SRE) interviews — common questions

Who is this Site Reliability Engineer (SRE) interview plan for?
It is written for the interviewer, not the candidate: the hiring manager, engineer or panel member running the Technical Interview round for a Site Reliability Engineer (SRE) role. It gives you a 60 min script to follow in the conversation — 4 questions with what each one is for and the signals to score against — so you are not writing the round from scratch the night before.
What does the Technical Interview round assess?
This round is focused on: SLOs and error budgets, alerting, and capacity planning, argued from one booking API's numbers: an SLO per journey, burn-rate pages tested against 139 real ones, a release on a two-thirds-spent budget, and sizing for a noon spike. It works through One number for three journeys, 139 pages and a phone call, Two-thirds of a budget, four days out and Twelve o'clock on Monday, scoring against 35 observable signals, with follow-up prompts on all 4 questions for going deeper where an answer is thin.
How is the 60 min split up?
60 min on 4 questions. The questions take in One number for three journeys (16 min), 139 pages and a phone call (16 min), Two-thirds of a budget, four days out (12 min) and Twelve o'clock on Monday (16 min). The timings are there so the round stays on schedule and every candidate gets the same shape of interview — which is what makes two candidates comparable afterwards.
What other rounds should I run for a Site Reliability Engineer (SRE)?

A single round does not cover a whole role. The other rounds in this library for a Site Reliability Engineer (SRE):

Hiring for this role?

Open this plan in Hirezen and make it a position in one click.

  • Every interviewer runs the same script and marks the same signals.
  • AI drafts the write-ups, and the debrief puts every read side by side.
  • No ATS to set up first, and no bot in the call.
Start free with this plan

Free for your first open role.