Interview plan template
A 60 min interview plan with a time-boxed script, what each question is for, and the signals to score against. Key skills: Whether the candidate can measure an LLM feature rather than describe one: partitioning real failure traces before proposing a fix, interrogating an LLM judge's own error rates, designing an instrument two annotators would agree on, knowing the noise floor of a shipping decision, and attaching money and an owner to the eval set..
I'm [YOUR_NAME] and I build the LLM features at [COMPANY_NAME]. This hour is about measurement, not architecture — I am not going to ask you to design a retrieval pipeline, I am going to hand you one that is already broken and watch how you find out why.
Here is the pack: six traces from a support-answering feature, sampled from a week when users kept hitting regenerate. Each one has the user's question, the chunks we retrieved with their scores, the answer we returned, and what the answer should have been. Take two minutes to read before you say anything — I would rather you were quiet than fast.
There is no trick in the pack and no single root cause — expect more than one thing to be wrong. If you think one of the six is not actually a failure, say so; that counts.
Separates candidates who diagnose from candidates who prescribe. The entire discriminator is whether a proposed fix arrives before a partition of the failures does.
Switching from the traces to the dashboard. Assume this feature has been in production a quarter, and there is a nightly eval job whose numbers get pasted into a channel every morning.
The most reliable separator in the round. People who have run a judge in production reach for its error rates unprompted; people who have read about judges list biases and stop there. Read the arithmetic in the model answer before you run this — you will be asked to defend it.
A deliberately bad instruction. The signal is what the candidate does with an instrument they have been handed, and whether their objection is grounded in annotator agreement or in taste.
The first half has a widely believed half-answer, which makes the follow-up impossible to bluff. The second half is the decision the folklore actually costs money on.
Tests whether the eval set is treated as a living artifact with an owner and a decay rate, and whether the candidate checks the denominator before theorising about the model.
Forces both halves of the trade into one answer. Cost work with no quality metric and eval work with no bill are each half an answer, and most candidates volunteer exactly one half. Three asks in nine minutes is already tight — do not add the fourth.
Last thing, and it is the part I weight most: what do you want to ask me about how we measure this feature? I will answer anything — who owns our eval set, whether our judge gates, what the review load looks like on a bad week.
To close the loop honestly: [name one true, unflattering fact about your own measurement — an unvalidated judge, an eval set with no owner, a dashboard nobody acts on]. If you take this job you inherit that. I would rather you knew it now than found it in week three.