Interview plan template
A 60 min interview plan with a time-boxed script, what each question is for, and the signals to score against. Key skills: Scoping an underspecified AI deployment inside a customer's own environment, under a live role-play with access constraints, a late scope change and a handback.
I am [YOUR_NAME] at [COMPANY_NAME], and for the next hour I am not going to be an interviewer. I am going to be a customer - I run operations at a company that has just signed a pilot with us, and you are the engineer we are sending in. I know a great deal about my business that I am not going to volunteer, because I do not know which parts of it matter to you. Ask me anything. If I do not know, I will say so, and that will be true. About twenty minutes in I am going to stop you and ask what will be running in my environment a week from now, so leave yourself room for that. There is no whiteboard requirement here - talk, and write down whatever you would want to send me afterwards.
So, here is where we are. I run operations at a mid-size commercial insurance broker. We get somewhere around four hundred emails a day from carriers and clients with documents attached - quotes, binders, endorsements, loss runs. Someone on my team opens each one, works out what it is, pulls the numbers off it, and types them into our policy administration system. I want your AI to do that. My CFO has already signed off on a pilot, and I would like to see something working this quarter.
Separates candidates by the order of their discovery, not by the presence of it. Every prepared candidate asks clarifying questions; the signal is whether the first ones are about the people and the clock or about the file formats.
Tests whether deployment reality is part of the plan or an afterthought. A candidate who has only ever built inside their own company's systems answers this in terms of architecture; a candidate who has deployed into someone else's answers it in terms of access, approvals and calendar days.
Distinguishes candidates who verify the data from candidates who trust the description of it. The brief describes a clean stream of PDFs; the reality underneath is dirtier and thinner, and discovering that in week four, and not in this hour, is what burns an engagement.
Let me stop you there. I have a hard stop at the top of the hour, and there is something I need from you before you go.
Reproduces the defining failure of the role inside one hour: the beautiful design nobody deployed. Candidates who sequence by architectural layer make nothing demonstrable until week three, and this question exposes that in ninety seconds.
Two things before you go. My CFO looked at this yesterday and asked whether it could also flag renewal quotes that come back more than ten percent above last year - that is the thing that actually loses us clients. And he does not want it live when renewal season starts. He wants it live four weeks before that, so my team has time to learn it before they are busy. That is seven weeks, not eleven.
The behaviour under scope pressure is the behaviour that decides whether engagements survive. Capitulation and flat refusal both feel decisive in the room and both fail; the pass is a priced trade with an explicit answer attached.
Nearly every AI-engineering guide teaches "build an eval harness", so the generic answer is universal and worthless. The discriminating move is borrowing the customer's own existing manual quality process as the first evaluation set, which only people who have shipped into a real business reach for.
The defining hazard of the role, and the part of the job that decides whether the sample ever arrives. The ledger already contains everything this question needs: the CFO is buying headcount cost, four of the six are agency contractors, and the supervisor whose spreadsheet the candidate has just asked for is one of the people being automated. Someone who has been embedded answers this instantly and specifically. Someone who has not talks about change management.
Tests the property that defines the role, which is leaving an artifact and not a recommendation, and simultaneously checks whether sixty minutes of good conversation produced anything shareable. Candidates who talk brilliantly and hand over nothing are demonstrating the exact behaviour that loses engagements. The readback is scored in the room; the written summary is scored later, against the post-round rubric in this section's opening note.
That is my hour, and I am stepping out of the customer role now. Send me the write-up today or tomorrow if you can - I will read it the way that customer would, which means I am looking for the out-of-scope list as closely as the plan. We will come back to you either way within three working days.