Interview plan template

Use Template
to edit & run interviews

Marketing Analytics Engineer interview questionsMeasurement Round — The Number Two Systems Disagree About round

A 60 min interview plan with a time-boxed script, what each question is for, and the signals to score against. Key skills: Decomposing a conversion discrepancy between an ad platform and the warehouse without assuming either is broken, separating attribution from incrementality, specifying a holdout somebody will actually act on, and refusing to build a metric that cannot be wrong.

Click "Use template" to edit

The discrepancy

10 min

I'm [YOUR_NAME] and I own the data platform at [COMPANY_NAME]. Two things about the hour. I'm going to give you a disagreement between two systems and I want to be clear up front that I do not think either of them is broken — if you go looking for a bug you may find one, but that is not where I expect this to land. And in the second half I'll ask you to specify an experiment precisely enough that we could run it, which is a different skill from describing one, so expect me to push on the details until they are decidable. If at any point you need a number I have not given you, ask for it; I will either have it or tell you honestly that nobody here knows, and which of those two it is will be useful to you.

Our ad platform reports 1,200 conversions last month. Our warehouse reports 780 for what everybody calls the same metric. Separately, if I add up what every platform claims across all our channels, the total is about 40% higher than the warehouse's total conversions for the month. Nobody has reconciled any of this in the eighteen months we have had the warehouse. Where do you start?

What this question is for, and what to listen for

Purpose

The opening discriminator, and it resolves on whether the candidate asks for definitions before proposing causes. There are at least eight ordinary reasons for this gap and none of them is a defect; a candidate who starts naming causes is guessing, and a candidate who starts asking what each system counts is doing the job.

Signals to score

  • Asks what event each system is counting, in precise terms, before proposing any explanation
  • Asks about the attribution window configured on each side, and knows the two are almost never the same
  • Asks whether the platform includes view-through conversions and the warehouse almost certainly does not
  • Names deduplication as the explanation for the cross-channel sum exceeding the total
  • Asks whether consent mode or any conversion modelling is enabled, and knows modelled conversions are estimates included in the reported figure
  • Asks about timezone and whether conversions are stamped at click time or conversion time
  • Asks how the warehouse joins a conversion to a channel at all, and what happens when the join fails
  • Raises browser storage limits and their effect on client-side identification
  • Treats the 40% cross-channel excess as a structurally different question from the 1,200-versus-780 gap
  • Proposes reconciling on a small, well-understood slice before attempting the whole month

Follow-up questions

  • Which of the two numbers is right?
  • Why does the sum across channels exceed the total? Isn't that just impossible?
  • Suppose the definitions turn out to match exactly. Then what?
  • How long would a first pass at this take you, honestly?
  • What would you need from me before you could start?

What the gap was made of

18 min

Say you decompose it and most of the gap resolves — windows, view-through, modelling, join failures. Fine. Now the question my CMO actually has. Paid brand search reports the best return of anything we run, by a distance. Every attribution model I've tried agrees. Should I keep funding it?

What this question is for, and what to listen for

Purpose

The pivot, and the point at which every attribution model in the room is unable to answer the question being asked. Branded paid search is the canonical case where attribution and incrementality diverge, because the conversions are real, the credit is arguably correct, and the spend may still be buying clicks that were arriving free. A candidate who answers from the attribution data has missed the entire discipline.

Signals to score

  • Says that no attribution model can answer this, and explains why rather than asserting it
  • Names the mechanism: someone searching our brand name was likely to arrive anyway, and the ad converts a free click into a paid one
  • States that every model agreeing is expected and is not corroboration, because they share the same observational limitation
  • Distinguishes attribution — dividing credit among observed touchpoints — from incrementality — what would have happened without the spend
  • Proposes an experiment as the only thing that settles it, rather than a better model
  • Names a specific design: geo holdout, switchback, or a staged pause, with matched controls
  • Measures total conversions rather than channel-attributed conversions, and says why that is the entire point
  • Raises the competitor-bidding confound, where pausing cedes the position to somebody else
  • Cites known evidence with a source and a date, or declines rather than inventing one
  • Estimates what the test costs in forgone conversions if the sceptical view is wrong

The experiment

20 min

Specify it. I want the design precise enough to hand to someone: which markets, how they're chosen, how long, what you measure, and the decision rule. And be specific about the decision rule — I want to know before we start what result means we cut the spend and what result means we keep it.

What this question is for, and what to listen for

Purpose

Pre-registering the decision rule is the single strongest predictor available of whether this person produces analysis that gets acted on. An experiment whose success criteria are agreed after the result arrives is not an experiment, and this organisation has already demonstrated it negotiates numbers.

Signals to score

  • Selects markets by pre-period similarity on the outcome, not by convenience or by size alone
  • Uses enough markets on each side that one anomaly cannot drive the result
  • States a duration and ties it to accumulating a decidable difference rather than to a calendar
  • Measures total conversions in the market as the primary outcome, and names it as primary before naming anything else
  • Pre-specifies the decision rule in both directions, with numbers
  • Acknowledges uncertainty in the estimate rather than treating the point difference as the answer
  • Names the confounds that would invalidate the run — a competitor's campaign, a promotion, a seasonal event in one market
  • Says what monitoring runs during the test and what would stop it early
  • Proposes a pre-period sanity check: run the comparison on historical data with no intervention and confirm it shows no effect
  • Says what they would do if the result is ambiguous, rather than assuming it will be clean

Last technical thing. Somebody is going to ask you to build a dashboard metric for AI search visibility — our brand's presence in assistant answers. Tell me how you'd build it, or tell me why you wouldn't.

What this question is for, and what to listen for

Purpose

Connects this role to the rest of the loop and tests whether the candidate will build a metric they know to be bad because somebody senior asked. The correct answer is not refusal; it is building it with the properties that make it capable of being wrong, and saying which questions it cannot answer.

Signals to score

  • Asks what decision the metric is meant to inform before designing it
  • Identifies non-determinism as the core measurement problem and requires repeated sampling
  • Reports variance or an interval alongside any central figure
  • Insists the prompt set be versioned and governed by someone who does not report the number
  • Names the failure mode where prompts naming the brand guarantee a mention
  • Requires a competitor baseline in the same measurement
  • Separates mention, recommendation and citation rather than summing them
  • Says plainly which questions the metric cannot answer, especially anything about revenue
  • Proposes shipping it alongside a caveat rather than refusing outright
  • Names a condition under which the metric should be retired

What you would not build

12 min

Two to finish. First: name a metric or a report you've been asked for that you'd push back on, and tell me how that conversation went — or how it would go here. Second: what would you need from me in the first ninety days for your numbers to be the ones people actually use?

What this question is for, and what to listen for

Purpose

The first half finds out whether the candidate has ever said no to a number, which is the defining act of this function. The second half is where the real blockers surface, and at this band they decide offers more often than compensation does.

Signals to score

  • Names a specific metric and a specific objection, not a general commitment to rigour
  • Recognises the class where the denominator is chosen by the person reporting the number
  • Offers an alternative rather than only an objection
  • Describes an actual conversation, including the part where they did not fully win
  • Distinguishes a metric that is imprecise from one that is structurally incapable of being wrong
  • Asks who currently owns the numbers people quote in meetings, and whether that is the warehouse
  • Asks what happens when their number contradicts a channel owner's number, and who adjudicates
  • Names a specific access needed on day one — the warehouse, platform configuration, the event definitions, the experiment tooling
  • Asks whether they can change a definition that people have built expectations on, and what that process looks like
  • Raises documentation or definition ownership as a deliverable rather than as hygiene

That's what I had. What I write up is your experiment design and the decision rule, and I'll be straight about what happens to it: I am going to take that design to the person who owns paid, and their reaction to a pre-registered decision rule is going to tell me something about my own organisation. I'll tell you what they say. Two things before you decide about us. Nobody has reconciled those numbers in eighteen months, which means the disagreement we spent the first section on is live and is the first thing you would own. And we have never run an incrementality test on anything, so if that interests you it is not a line in a job description here, it is genuinely unclaimed. [RECRUITER_NAME] will come back to you within [NUMBER] working days.

Use Template
to edit & run interviews
Interview Template
Position
Marketing Analytics Engineer
Round
Measurement Round — The Number Two Systems Disagree About for 60 min
Key skills
Decomposing a conversion discrepancy between an ad platform and the warehouse without assuming either is broken, separating attribution from incrementality, specifying a holdout somebody will actually act on, and refusing to build a metric that cannot be wrong

Marketing Analytics Engineer interviews — common questions

Who is this Marketing Analytics Engineer interview plan for?
It is written for the interviewer, not the candidate: the hiring manager, engineer or panel member running the Measurement Round — The Number Two Systems Disagree About round for a Marketing Analytics Engineer role. It gives you a 60 min script to follow in the conversation — 5 questions with what each one is for and the signals to score against — so you are not writing the round from scratch the night before.
What does the Measurement Round — The Number Two Systems Disagree About round assess?
This round is focused on: Decomposing a conversion discrepancy between an ad platform and the warehouse without assuming either is broken, separating attribution from incrementality, specifying a holdout somebody will actually act on, and refusing to build a metric that cannot be wrong. It works through The discrepancy, What the gap was made of, The experiment and What you would not build, scoring against 50 observable signals, with follow-up prompts on 1 of the 5 questions for going deeper where an answer is thin.
How is the 60 min split up?
The discrepancy (10 min), What the gap was made of (18 min), The experiment (20 min), What you would not build (12 min). The timings are there so the round stays on schedule and every candidate gets the same shape of interview — which is what makes two candidates comparable afterwards.