Hiring for this role?
Start free with this planFree for your first open role.
Why Hirezen?
- Every interviewer runs the same script and marks the same signals.
- AI drafts the write-ups, and the debrief puts every read side by side.
- No ATS to set up first, and no bot in the call.
Cloud Architect interview questionsSystem Design Interview round
A 60 min interview plan with a time-boxed script, what each question is for, and the signals to score against. Key skills: Multi-region disaster recovery design to a stated RTO and RPO: tiering a platform by what each part can afford to lose, finding what still depends on the region that failed, keeping EU-plan data in the EU down to backups and logs, and writing the recovery numbers a board actually signs.
Opening
Who is interviewing, how the round will run, and a question to settle the candidate in. The standard openingWhat the board asked for
What this part is for
Purpose
Runs over a design pack sent 24 hours ahead. Write it once and give every candidate the same one. The company is a B2B invoicing and payments platform with 3,000 business customers, in one cloud region — Region A, in the EU — across three availability zones. The pack holds six things. A component table: a stateless API and web tier in containers on a managed Kubernetes cluster; an invoice database on managed PostgreSQL, 2.4 TB, with a synchronous standby in a second zone, peaking at 1,800 write transactions a second, where `create invoice` commits six transactions in sequence and runs at 180 ms p95 against a 300 ms budget; 38 TB of attachments in object storage; outbound webhooks on a managed queue; a cache of rendered dashboards and rate-limit counters; and card payments through an external processor whose API lists every event from the last 30 days. A backup note: a nightly snapshot copied to Region B, encrypted with a key held in the company's own key manager, which runs only in Region A; transaction logs archived every five minutes to object storage in Region A only, and a last full restore test, in Region A, that took 3 hours 40 minutes. An operations note: container images live in a registry in Region A only, and engineers reach the cloud accounts only through single sign-on configured in Region A. A latency and cost sheet: Region B, also in the EU, is 24 ms round trip at p50 and 31 ms at p99; Region C, in the US, is 92 ms and prices 12% lower; a test cross-region replica into Region B lagged 2.5 seconds at p99 at peak; Region A costs $62,000 a month, and the team estimates Region B at +$2,400 for backup and restore, +$21,000 for a warm standby and about +$64,000 for active-active. A contract clause: customers on the EU plan, 31% of revenue, are promised that their data, including backups, is stored and processed only in the EU — and a line saying application logs, which include request bodies for failed API calls, ship to an observability vendor's US endpoint. And the board's request, after a provider regional outage last year took the platform down for five hours: survive the loss of a region "with no data loss and no more than 15 minutes of downtime", with the platform team's draft answer — "warm standby in Region C; it is cheaper and the analytics warehouse is already there". The stateless tier is deliberately fine as it is, and the cache is the red herring: it rebuilds cold in minutes, and a candidate who spends time replicating it is designing by habit. Book 70 minutes; the last exchange sits outside the 60.
I'm [YOUR_NAME] and I look after platform architecture at [COMPANY_NAME]. You have had the pack since yesterday. This is a design hour, but the platform already exists and already has a board asking for something — I care less about how many boxes you draw than about which numbers you are prepared to promise.
What this line is for
Purpose
Frames the round as a design against a real request and a real estate, so a candidate who prepared a generic high-availability diagram knows in the first minute that it will not be enough.
The board's words are in the pack exactly as they said them. Treat them as the start of a conversation with the board, not as a specification you are bound by.
What this line is for
Purpose
Gives permission to negotiate the requirement. Whether the candidate uses it, and how early, is the first thing this section reads.
The board wants to survive losing Region A with no data loss and fifteen minutes of downtime. Using the pack, tell me what you would build for each component, and what recovery point and recovery time each one actually gets.
What this question is for, and what to listen for
Purpose
The load-bearing question. Any candidate can draw two regions; the separator is whether the design is set component by component against numbers from the pack, and whether the candidate finds out what "no data loss" costs before promising it.
Signals to score
- Sets a recovery point and recovery time per component instead of accepting one target for the whole platform
- Works out what synchronous replication to Region B does to `create invoice`: six commits each waiting 24 to 31 ms, well past the 300 ms budget
- Says that a synchronous commit across regions ties Region A's availability to Region B's, so "no data loss" costs latency and uptime or is not true
- Turns the 2.5 second replica lag into a number the board can weigh — about 4,500 writes at peak, several hundred invoices
- Uses the payment processor's 30-day event history to recover card payments after a failover instead of chasing zero lag in the database
- Rejects backup and restore for the invoice database on the 3 h 40 min restore and on transaction logs that never leave Region A
- Chooses a warm standby with a cross-region replica that is promoted on failover, and says what already runs in Region B before the day
- Leaves the cache out of replication and says why
- Picks Region B over Region C and gives the EU plan clause as the reason, not latency
- Puts a monthly cost on the recommendation from the pack's own estimates
Follow-up questions
- The board said no data loss. What would it take to make that literally true, and what would it do to `create invoice`?
- The replica lags 2.5 seconds at peak. How much is that in invoices, and whose are they?
- Card payments settle through the processor. Does that change what "no data loss" has to mean for money?
- Region C is 12% cheaper and the analytics warehouse is already there. Why not use it?
- What in this estate would you not replicate at all?
What still lives in Region A
What this part is for
Purpose
A recovery region is only as good as the list of things it does not need from the region that failed. The pack plants three such dependencies — the image registry, the key manager holding the key for the snapshot copies, and single sign-on — and labels none of them. The transaction logs archived only to Region A are a fourth, already met in the first section.
Region A has been unreachable for twenty minutes and you are failing over to your design in Region B. Walk me through what Region B needs in its first ten minutes, and tell me which of those things the pack says exist only in Region A.
What this question is for, and what to listen for
Purpose
Tests whether the candidate designs recovery as a sequence of dependencies rather than as a second copy of the diagram. Candidates who have run a failover go straight to access and images; candidates who have drawn one go straight to the database.
Signals to score
- Finds that container images exist only in Region A's registry, and replicates them to Region B ahead of time
- Finds that the snapshot copies in Region B are encrypted with a key whose only key manager runs in Region A, and says they cannot be restored while Region A is gone
- Finds that engineers reach the accounts only through single sign-on configured in Region A, and proposes emergency access that does not depend on it, with how it is held and how its use is noticed
- Notes that transaction logs archive only to Region A, so the backup path in Region B recovers to last night, not to five minutes ago
- Checks that Region B has the capacity and quota to run full production, not only the standby's footprint
- Separates what must already be running in Region B from what can be created during failover, and times the second list against fifteen minutes
- Asks what else is regional that the pack does not mention — secrets, certificates, the infrastructure code's state — instead of stopping at three
- Says what kind of exercise would have found each dependency, and that only a real failover finds all of them
Follow-up questions
- Where does Region B pull its container images from on the day?
- The snapshot copies are already in Region B. Can you restore one?
- How does the first engineer on the call get into the account?
- Which of these would a walk-through on paper find, and which only a real failover?
- What else is regional that is not in the pack?
Where the EU plan's data goes
What this part is for
Purpose
Residency is a boundary drawn around data, not a choice of region, and the pack has two places where it is already crossed or about to be: the draft answer's Region C, and the application logs. Neither is flagged as a residency problem in the pack.
Thirty-one per cent of revenue is on the EU plan. Trace one EU-plan customer's invoice through your design — database, attachments, backups, logs, support access — and tell me everywhere it is stored or processed.
What this question is for, and what to listen for
Purpose
Tests whether the candidate follows data past the architecture diagram into backups, logs and people, and whether they treat a boundary already crossed as today's problem rather than as a note for the recovery project.
Signals to score
- Rules out Region C for EU-plan data on the clause alone, before any argument about cost
- Finds that request bodies from failed API calls reach a US logging endpoint, and treats it as a problem now, not as part of the recovery design
- Includes snapshot copies and archived logs in the trace, since the clause names backups
- Asks what "processed" means in the contract — whether support staff outside the EU opening an invoice counts — and says who should answer that
- Proposes a fix at the boundary for logs — strip request bodies before they leave, or ship to an endpoint in the EU — and says which comes first
- Weighs splitting recovery by plan — EU-plan customers to Region B, the rest to Region C — against the roughly $2,500 a month it would save
- Asks where the keys that encrypt EU-plan data are held and whether that matters under the clause
- Names who hears about the logging finding, and before whom
Follow-up questions
- The logging vendor's endpoint is in the US. Is that a recovery design problem?
- Could EU-plan customers recover to Region B and everyone else to Region C?
- A support engineer outside the EU opens an EU-plan customer's invoice. Is the clause broken?
- What do you do about the logs this week, before any recovery work starts?
- Who do you tell, and what do you say?
The paragraph the board signs
What this part is for
Purpose
The design is only a decision once the people who asked for something else have agreed to what they are getting. The candidate writes the short text a board can sign and hold someone to.
The board asked for zero data loss and fifteen minutes. Your design gives them something else. Write the paragraph they sign — what they get, what it costs, and what it does not protect them from — and read it to me.
What this question is for, and what to listen for
Purpose
Tests whether the candidate can turn a design into a commitment in plain numbers, including the part the board did not ask about. A paragraph full of reassurance is the answer this question exists to catch.
Signals to score
- States recovery point and recovery time in numbers, per tier where they differ, instead of "near-zero" or "minimal"
- Puts the recovery point in business terms — writes at risk at the worst moment, and that card payments are recovered from the processor's records
- States the monthly cost, and what the cheaper and the dearer options would have bought
- Says what the design does not cover — a bad release or a deleted table reaches Region B within seconds — and what does cover it
- Names the decision the board is taking and who owns the recovery numbers once it is taken
- Says how the numbers will be proven, and what the board will see if a test misses them
- Keeps it short enough to read aloud in about a minute, with no product names
- Offers the path to what the board asked for, with its price, instead of declaring it impossible
Follow-up questions
- A board member asks why not zero. Answer in one sentence.
- Someone runs a migration that deletes a table in production. What does your design do about it?
- Which number in your paragraph are you least sure of, and how will you find out?
- Who owns the fifteen minutes once the board has signed?
Closing
We have a few minutes left. What do you want to know about how recovery really works here? Ask anything — when we last failed over, who owns the recovery numbers, what our last restore test found.
What this line is for
Purpose
A candidate who has owned a recovery target asks when it was last proven and by whom; one who has not asks which provider we use. Offering topics makes what they choose the signal.
One honest thing before you go: [name one true gap in your own recovery setup — a restore nobody has timed, a key or a registry that lives in one region, a recovery target nobody has signed]. Whoever takes this job starts there.
What this line is for
Purpose
A specific, unflattering fact is the pitch that works on the candidate this round is looking for, and it puts off one who wants the recovery work already done. Check that it is still true before you say it.
Cloud Architect interviews — common questions
- Who is this Cloud Architect interview plan for?
- It is written for the interviewer, not the candidate: the hiring manager, engineer or panel member running the System Design Interview round for a Cloud Architect role. It gives you a 60 min script to follow in the conversation — 4 questions with what each one is for and the signals to score against — so you are not writing the round from scratch the night before.
- What does the System Design Interview round assess?
- This round is focused on: Multi-region disaster recovery design to a stated RTO and RPO: tiering a platform by what each part can afford to lose, finding what still depends on the region that failed, keeping EU-plan data in the EU down to backups and logs, and writing the recovery numbers a board actually signs. It works through What the board asked for, What still lives in Region A, Where the EU plan's data goes and The paragraph the board signs, scoring against 34 observable signals, with follow-up prompts on all 4 questions for going deeper where an answer is thin.
- How is the 60 min split up?
- 60 min on 4 questions. The questions take in What the board asked for (24 min), What still lives in Region A (12 min), Where the EU plan's data goes (13 min) and The paragraph the board signs (11 min). The timings are there so the round stays on schedule and every candidate gets the same shape of interview — which is what makes two candidates comparable afterwards.
- What other rounds should I run for a Cloud Architect?
A single round does not cover a whole role. The other rounds in this library for a Cloud Architect:
Hiring for this role?
Open this plan in Hirezen and make it a position in one click.
- Every interviewer runs the same script and marks the same signals.
- AI drafts the write-ups, and the debrief puts every read side by side.
- No ATS to set up first, and no bot in the call.
Free for your first open role.