Interview plan template

Use Template
to edit & run interviews

AI Search & SEO Manager (AEO/GEO/AIO) interview questionsStrategy Round — What the Answer Engine Said round

A 60 min interview plan with a time-boxed script, what each question is for, and the signals to score against. Key skills: Measuring brand presence in a non-deterministic answer engine, separating a zero-click story from a competitive loss before spending against either, auditing a visibility metric that was satisfied by its own failure mode, and deciding which surfaces are worth defending once the answer arrives without the click.

Click "Use template" to edit

Framing and the first measurement

10 min

I'm [YOUR_NAME], I own search at [COMPANY_NAME], and this round is the strategy half — there's a separate technical round with an engineer, so you and I are not going to spend the hour on rendering or status codes. Three things about the format. I'll put real numbers from our business in front of you, including a couple I'm not proud of, and I would rather you interrogate them than accept them. At two points I'll ask you to design a measurement rather than describe a strategy, and by the end I want one metric definition precise enough that I could hand it to an analyst on Monday. And I will tell you where I currently think the answer is, at least once, specifically so you can disagree with me — the last person who took my framing at face value for a full hour is not the reason this round exists, but they are the reason it changed. Being wrong out loud is fine here. Being agreeable is the failure mode I'm actually screening for.

Here's a browser. The query our highest-intent buyer actually types is on the table — for us it's "best software for booking salon appointments". Ask it of ChatGPT, Perplexity and Google's AI Overview. Run each one three times. You've got four minutes for the running, and then I want your read. But before you touch anything: tell me what you're measuring, and what you expect the three runs to do.

What this question is for, and what to listen for

Purpose

The fastest discriminator available and it resolves before any result loads. Someone who has actually run an AEO programme knows the three runs will disagree and says so up front, which reframes the exercise from observation to sampling. Someone who has read about one screenshots the first answer and reads it like a SERP. The pre-commitment is the whole point — asking afterwards lets a quick candidate reverse-engineer the insight from what they see.

Signals to score

  • States before running that the three runs will not match, and that a single observation is a sample rather than a position
  • Asks for a signed-out or fresh session without being prompted, and names memory, personalisation or account history as the contaminant
  • Distinguishes what they are counting: mention, recommendation, and citation with a resolvable link are three different events
  • Asks who the competitors in the answer are, and treats the absence of a baseline as the thing that makes any number meaningless
  • Notices whether the answer names a source it did not link, or links a source it did not name
  • Reads position within the answer — first named, listed among eight, mentioned in a caveat — rather than scoring presence as binary
  • Names sentiment as a dimension: "cheaper but limited" is a loss that a presence metric records as a win
  • Asks what the model can even know, and whether the grounding is retrieval at query time or training data with a cutoff
  • Proposes a repeatable protocol — prompt set, runs per prompt, cadence, who runs it — rather than a one-off audit
  • Says out loud that four minutes buys an anecdote and names what sample size would buy a number

Follow-up questions

  • You ran it three times and got three different answers. So is the metric useless?
  • Which of the three engines do you trust least, and why?
  • Give me the smallest protocol that would actually tell me something. Prompts, runs, cadence.
  • We're mentioned in all three. Good day?
  • If I gave you a vendor tool that tracks this automatically, what would you still want to check by hand?

The twenty-two per cent

18 min

Here are three numbers from the same twelve months. B2B organic sessions are down 22% year over year. Our brand appears in AI Overviews for roughly three times as many queries as it did last year — that's from our own tracking, and you can push on it. Demo requests are flat. Our CMO has read all this as a zero-click story: we're winning the answer and losing the click, demand is intact, and we should stop worrying about sessions. Three different stories fit those numbers. Name them. Tell me which you'd bet on and at what confidence. Then give me the single measurement that separates them, and tell me what result would kill your own answer.

What this question is for, and what to listen for

Purpose

Discriminates between a candidate who can hold competing explanations and one who adopts the executive's framing because it is already in the room and it is flattering. The zero-click story is plausible, it is what the CMO believes, and it is the one this round is built to make somebody argue with. The tell is the order: whether the dull explanation gets ruled out before the interesting one gets adopted.

Signals to score

  • Names the zero-click story, a competitive-loss story, and at least one measurement-artifact story, before ranking any of them
  • Reaches the boring explanation unprompted: a tracking or consent change, a reporting definition change, a migration, a core update, seasonality, a bot-filtering change
  • Refuses to treat "organic sessions" as one number and immediately asks to split by query intent
  • Points out that our AI Overview tracking is our own and asks how the prompt set was built before accepting the 3× figure
  • Names the specific separator: informational versus commercial-intent queries compared separately in Search Console, impressions against clicks against CTR
  • States what a genuine zero-click pattern looks like — impressions flat or up, CTR falling, concentrated on informational queries — and what a competitive loss looks like instead
  • Knows that Search Console does not break out AI Overview impressions as a separate dimension, and says what that costs the analysis
  • Reads "demo requests are flat" as ambiguous rather than reassuring, and asks what happened to demo request rate per session
  • Gives a confidence number when asked and it is not 90%
  • Names the result that would falsify their own pick, without being asked twice

Follow-up questions

  • Demo requests are flat. Doesn't that settle it?
  • Our AI Overview tracking says 3×. How much weight do you put on that number?
  • Suppose commercial-intent impressions are also down. What now?
  • Give me a confidence number, and tell me what would move it.
  • The CMO wants to reallocate the content budget to AI optimisation next quarter based on this. What do you tell them?

The metric that went up while we lost

20 min

My predecessor built our AI search dashboard. It reports one headline number — "AI Visibility Score", defined as the percentage of tracked prompts where our brand is mentioned in the answer. It went from 34% to 61% over two quarters. In those same two quarters we lost the category: the competitor everybody now names as the default was a distant third when that dashboard was built. The number is not fabricated. Nobody lied. Tell me why it went up.

What this question is for, and what to listen for

Purpose

Anybody can agree that a better metric is needed; this asks why the metric that already existed was worthless, and the generalisation the candidate draws predicts every instrument they will build for us afterwards. The intended realisation is that the metric was satisfied by the failure mode, which is the same defect class the engineering round finds in a CI assertion, in a different medium.

Signals to score

  • Asks who wrote the tracked prompt set, and when it was last changed, inside the first two minutes
  • Identifies that prompts naming the brand guarantee a mention, making the metric partly a measurement of its own prompt list
  • Separates mention from recommendation from citation, and says the metric collapses all three
  • Names the missing competitor baseline: 61% is meaningless without knowing the leader is at 95%
  • Points out that position within the answer and sentiment are uncaptured, so a mention as the cheap-but-limited option scores as a win
  • Asks how many runs per prompt, and treats a single run per prompt as noise recorded as signal
  • Notes that a prompt set that grows over time produces a rising percentage with no change in reality
  • Asks whether the prompt set was ever refreshed against how the market actually searches now
  • Generalises to the class — the metric could not have fallen while the thing it measured got worse — without being led there
  • Declines to blame the predecessor, and treats a metric that cannot go down as the reportable finding

Follow-up questions

  • Would a stricter definition have saved us? Say, mentioned in the first three sentences?
  • If you'd inherited this dashboard, what would have made you suspicious of it on day one?
  • We also tracked branded search volume, and that rose too. Why didn't that help?
  • What's the general version of this mistake, outside search entirely?
  • Where else in a marketing reporting stack would you go looking for the same defect?

Then replace it. Define the metric you'd put on that dashboard instead — precisely enough that I could hand the definition to an analyst on Monday and they'd build the same thing you have in your head. I want the numerator, the denominator, where the prompt set comes from and who owns it, the sampling, and what gets reported alongside the headline figure. Then two things. Tell me the number you'd expect it to read in month one, and tell me the condition under which you'd retire the metric you've just designed.

What this question is for, and what to listen for

Purpose

The artifact, and the only point in this round where the candidate produces something rather than critiques something. The specifics — how the prompt set is governed, what sits beside the headline number, the month-one expectation — are not improvisable from general knowledge, and the retirement condition is the single best predictor available of whether this person maintains instruments or accumulates them.

Signals to score

  • Gives a numerator that is narrower than mention: recommended, or cited with a resolvable link, and says which and why
  • Gives a denominator that is a governed prompt set with a stated origin — sales call language, Search Console commercial-intent queries, support tickets, win/loss notes — rather than one written by the search team alone
  • Freezes or versions the prompt set, so the number cannot move because the list moved, and says how a change is logged
  • Specifies runs per prompt and reports variance or a confidence interval next to the mean
  • Puts competitor presence in the same measurement, so the headline number always ships with a baseline
  • Separates the surfaces rather than averaging across engines that behave differently
  • Reports position or sentiment alongside presence, or explicitly defers it and says why
  • Names a month-one expectation and it is low, with a reason, rather than aspirational
  • Names a retirement condition tied to the metric ceasing to discriminate, not to the metric being unpopular
  • Volunteers what the metric will not tell us, unprompted

Follow-up questions

  • Who writes the prompt set, and what stops it drifting to make us look good?
  • I'll run your metric against last year's data. Does it show the loss we actually took?
  • The number comes back at 12% in month one. My CEO sees that slide. What happens?
  • How often does this get reported, and what decision does it change?
  • Name the thing your metric still cannot see.

What you would concede

12 min

Two surfaces. The B2B side, where somebody researches booking software and we want a demo request. And the marketplace, where a consumer looks for a provider near them and books an appointment. Assume an assistant can answer both queries well, and sends no click either time. Which of those do you fight, which do you concede, and tell me what fighting actually means in practice. Then the second thing, and I'd rather you spent real time here: what would you need from us in your first ninety days for any of this to be more than a slide?

What this question is for, and what to listen for

Purpose

"Fight both" is the absence of a position dressed as ambition, and the first half catches it. The surfaces have genuinely different economics and a candidate who has run search inside a business with a transaction in it will say so. The second half is a real question — someone who has owned this work has been blocked by data access, engineering capacity or ownership, and asks about whichever one burned them.

Signals to score

  • Gives different answers for the two surfaces and grounds the difference in the transaction, not in a preference
  • Says the click is the product on the marketplace side, because a booking cannot complete inside the answer
  • Names what an assistant cannot reproduce — live availability, current price, the booking itself — as the marketplace defence
  • Treats a B2B citation as carrying value without a click, because it shapes a shortlist that gets acted on later
  • Says the two need separate metrics, and that rolling them into one organic number is how you lose one while celebrating the other
  • Raises the possibility of conceding a surface deliberately and reallocating, rather than defending everything
  • Asks whether we want to be in the answer at all on surfaces where the assistant transacts, and treats crawler access as a decision with a cost either way
  • Asks who owns the marketplace pages and whether search changes there enter somebody else's roadmap
  • Names one specific access they want in the first ninety days — the log store, Search Console, the data warehouse, the CMS, the repository — rather than asking about tooling in general
  • Asks what happens the first time their recommendation costs another team a metric they are measured on

Follow-up questions

  • Concede the marketplace? Say that to the CEO.
  • What would you want to publish that a model can't produce for itself?
  • Would you block the AI crawlers on the marketplace side? Argue both directions.
  • Ninety days in, what's the one thing you'd want in writing?
  • Your recommendation costs the paid team their best-converting campaign. Who decides?

That's what I had. What I write up is the metric definition you gave me, and I'm going to do something with it: I still have the old dashboard and last year's data, so I'll run your definition against the two quarters we lost and tell you whether it would have caught it — including if it wouldn't. You'll hear that either way. Two things you should know before you decide anything about us. The 22% is unresolved as I speak to you, and whoever takes this job inherits that question rather than a conclusion about it. And the dashboard we spent twenty minutes on is still the one that goes to the board every month, which is the actual first project and the reason this round exists in this shape. [RECRUITER_NAME] will come back to you within [NUMBER] working days.

Use Template
to edit & run interviews
Interview Template
Position
AI Search & SEO Manager (AEO/GEO/AIO)
Round
Strategy Round — What the Answer Engine Said for 60 min
Key skills
Measuring brand presence in a non-deterministic answer engine, separating a zero-click story from a competitive loss before spending against either, auditing a visibility metric that was satisfied by its own failure mode, and deciding which surfaces are worth defending once the answer arrives without the click

AI Search & SEO Manager (AEO/GEO/AIO) interviews — common questions

Who is this AI Search & SEO Manager (AEO/GEO/AIO) interview plan for?
It is written for the interviewer, not the candidate: the hiring manager, engineer or panel member running the Strategy Round — What the Answer Engine Said round for a AI Search & SEO Manager (AEO/GEO/AIO) role. It gives you a 60 min script to follow in the conversation — 5 questions with what each one is for and the signals to score against — so you are not writing the round from scratch the night before.
What does the Strategy Round — What the Answer Engine Said round assess?
This round is focused on: Measuring brand presence in a non-deterministic answer engine, separating a zero-click story from a competitive loss before spending against either, auditing a visibility metric that was satisfied by its own failure mode, and deciding which surfaces are worth defending once the answer arrives without the click. It works through Framing and the first measurement, The twenty-two per cent, The metric that went up while we lost and What you would concede, scoring against 50 observable signals, with follow-up prompts on all 5 questions for going deeper where an answer is thin.
How is the 60 min split up?
Framing and the first measurement (10 min), The twenty-two per cent (18 min), The metric that went up while we lost (20 min), What you would concede (12 min). The timings are there so the round stays on schedule and every candidate gets the same shape of interview — which is what makes two candidates comparable afterwards.
What other rounds should I run for a AI Search & SEO Manager (AEO/GEO/AIO)?

A single round does not cover a whole role. The other rounds in this library for a AI Search & SEO Manager (AEO/GEO/AIO):