Hiring for this role?

Start free with this plan

Free for your first open role.

Why Hirezen?
  • Every interviewer runs the same script and marks the same signals.
  • AI drafts the write-ups, and the debrief puts every read side by side.
  • No ATS to set up first, and no bot in the call.

GTM Engineer interview questionsWork sample — the outbound workflow's first month round

A 60 min interview plan with a time-boxed script, what each question is for, and the signals to score against. Key skills: Whether the candidate reads an automated outbound workflow by what it produced rather than what it reports: taking a reply rate apart, finding out how much pipeline would have happened anyway, tracing a wrong personalised line to the data behind it, and deciding what to stop on the first morning they own it.

Opening

Who is interviewing, how the round will run, and a question to settle the candidate in. The standard opening

Reading the Pack

16 min
What this part is for

Purpose

A work sample, not a build. The candidate gets a one-page pack at the start and eight minutes to read it, and the round is spent on what they make of it. Print it or share it as a document; do not send it ahead, because what is read here is what the candidate notices on a first pass. ADAPTING THIS ROUND. This plan is published, so a candidate may have read it. Change the numbers before you use it and keep their shape: a reply rate swollen by automatic replies, more removal requests than interested replies, opportunities on accounts already in conversation, contacts at customers, one sending domain being refused, and one wrong personalised line. If your own outbound has a month like this, use it instead. THE PACK, in this order. // The workflow, built by a contractor who has since left: every morning it checks 2,000 target accounts for three triggers - a new head of the buying team in the last ninety days, a funding announcement in the last sixty, or a job posting that mentions our category. For each account that fires, an enrichment vendor finds and verifies up to four contacts, a language model writes an opening line from each contact's profile and the company's recent news, and the contact is enrolled in a four-step email sequence sent from thirty mailboxes on three secondary domains. // September: 4,200 contacts enrolled at 1,310 accounts. Step-one bounce rate 5.8%. The sequencer reports 382 replies, a reply rate of 9.1%, against 2.4% for the old hand-built sequence last spring. // The sales development reps (SDRs) tagged every reply: 104 automatic replies, 88 asking to be removed, 121 not interested, 23 wrong person, 46 interested. // 41 meetings booked, 26 held, 9 accepted by an account executive (AE) as qualified, 7 opportunities, $310,000 in pipeline, as posted in the sales channel. // 63 contacts were at existing customers, and 9 more at other accounts with an open opportunity. // In the third week, receiving mail servers began refusing one of the three sending domains: 14% of its step-one emails bounced that week, most citing the receiver's spam policy. // A reply a CTO sent, forwarded by the AE who owns his account: the opening line congratulated him on a Series B his company had not raised. // The SDR manager in the sales channel: "Best month we have had." The AE team lead: "The meetings are bad." IF THEY ASK -> YOU SAY, so every candidate gets the same answers. // What the workflow was for -> Qualified meetings for the AEs. The contractor's brief said twenty a month. // How the old sequence counted replies -> The same way, automatic replies included. Nobody tagged its replies. Its list was 900 contacts picked by hand. // Why so many accounts fired in one month -> A job posting stays up for weeks and the check runs daily, so an account keeps firing until the posting comes down. About 500 of September's accounts were also emailed in August, and 700 of the contacts had been enrolled before. // The seven opportunities -> Three are on accounts an AE had emailed or called in the sixty days before. Those three are $190,000 of the $310,000. // How long an opportunity takes to appear -> Three to eight weeks from a first email, judging by the CRM's dates. // What happens when someone asks to be removed -> The SDR marks them in the sequencer, which stops that sequence. Nothing else is told: not the other mailboxes, not the CRM. // How existing customers are excluded -> The workflow checks the CRM for the contact's email address. It does not check the account. // How the model finds company news -> A web search on the company name. Nothing checks that the article is about the same company. // Where the Series B came from -> An article about a different company with a similar name. // Sending volume -> About 650 emails a business day across the thirty mailboxes, split evenly across the three domains. // What it costs -> About $4,000 a month in tools. // Where the contacts are -> Mostly the US, some in the UK. // Who decides what the rules for cold email are -> Our general counsel, who has not been asked.

•

I am [YOUR_NAME] and I run go-to-market systems at [COMPANY_NAME]. This is a workflow a contractor built for us. It has been running for four months, the contractor has gone, and if you join, it is yours. Here is its September. Take eight minutes with it, and then I am going to ask you whether it works. Ask me anything the page does not say.

What this line is for

Purpose

Says the workflow will be the candidate's, which is true of most GTM engineering hires and turns the round from critique into ownership.

•

Is this workflow working?

What this question is for, and what to listen for

Purpose

Commercial judgment is read here. Includes the eight minutes of reading. The pack is built so that every headline number is defensible and none of them answers the question. The read is whether the candidate takes the reply rate apart, follows the funnel to something the company is paid for, and asks what would have happened without it.

Signals to score

  • Asks what the workflow was supposed to produce before judging whether it does
  • Takes the 104 automatic replies out of the reply count, and the removal requests out of what counts as success
  • Works the funnel down to qualified meetings and opportunities, and gives those as rates per thousand contacts
  • Asks whether the old sequence's 2.4% was counted the same way and sent to a comparable list before comparing
  • Asks whether the opportunities were on accounts already in conversation with a rep
  • Asks why 1,310 of 2,000 accounts fired in one month, and finds the same accounts and people being emailed again
  • Counts the costs in the pack - removal requests, the refused domain, customers contacted - alongside the results
  • Notices that the AE lead's complaint and the 9 qualified out of 26 held are the same finding
  • Gives a verdict with the conditions under which it would change, rather than declining to give one
  • Separates what the pack shows from what it would need to know

Follow-up questions

  • What is the reply rate once the automatic replies are out?
  • Which number in the pack would you report to the head of sales?
  • Is 9.1% against 2.4% a fair comparison?
  • Why might the AEs and the SDR manager both be right?
  • What would you need to know to say yes or no with confidence?

Would It Have Happened Anyway

14 min
What this part is for

Purpose

The core of the round. Every outbound system takes credit for pipeline it touched, and the GTM engineer is usually the only person placed to find out how much it caused. This is the section where the candidate designs that.

•

Our chief revenue officer has read the sales channel and wants to triple the volume in October.

What this line is for

Purpose

The pressure is real and comes from the most senior person in the room. The candidate is not asked to refuse it.

•

Before anyone triples it, how would you find out how much of that pipeline the workflow created? I want an answer before the end of the quarter.

What this question is for, and what to listen for

Purpose

Commercial judgment is read here. Reads whether the candidate knows that only a comparison against accounts the workflow did not touch can answer this, and whether they can design one that a sales team will tolerate and a revenue leader will accept.

Signals to score

  • Proposes holding out a random share of target accounts from the workflow, and randomises by account rather than by contact
  • Keeps reps free to work held-out accounts as they normally would, and says why that matters
  • Measures opportunities and pipeline per account in each group, not replies or meetings
  • Picks a duration from how long an opportunity takes to appear after first contact, and says what the test cannot see, such as closed revenue
  • Says that at about seven opportunities a month, a test this short can show only a large difference, and agrees that before it starts
  • Agrees before the test starts what result means scale, keep or stop
  • Uses the data already there first, such as prior rep activity on the seven opportunities' accounts, and says what it cannot prove
  • Offers more volume inside the test, so the request is answered rather than blocked
  • Counts the cost per qualified meeting, removal requests and refused mail included, alongside the pipeline

Follow-up questions

  • Why by account and not by contact?
  • What happens if a rep emails a held-out account anyway?
  • How long does the test need to run, and why?
  • What result would make you tell the revenue officer to triple it?
  • What do you tell them this week, before the test has an answer?

The Line the Model Wrote

14 min
What this part is for

Purpose

The model-written opening line is what makes this workflow look modern, and it is where a data problem became a message to a CTO. The read is whether the candidate finds the data failure under the model failure.

•

Read the CTO's reply again. How did the workflow come to write that line, and what changes before the next send?

What this question is for, and what to listen for

Purpose

Data quality is read here. A candidate can blame the model in one sentence. The mechanism is that the news was found by company name and never checked against the company, so the model faithfully used a fact about someone else. That is a matching failure, and the fixes are about data before they are about prompts.

Signals to score

  • Asks how the company news was found before blaming the model
  • Finds that the news was matched by company name and never checked against the company's domain or another identifier
  • Treats it as a matching problem first and a model problem second
  • Allows the line to use only facts that came from a source tied to the account, and writes no claim when there is none
  • Proposes reading a sample of generated lines and counting how many are wrong before the next send
  • Keeps the inputs beside each generated line, so a complaint can be traced to its source
  • Sets the acceptable error rate by who receives the email, not as one number for everyone
  • Pauses personalisation for contacts whose company match is uncertain, rather than pausing the whole workflow

Follow-up questions

  • Where did the Series B come from?
  • What would you check before blaming the model?
  • How many other lines this month were wrong, and how would you find out?
  • What does the line say when there is no news you trust?
  • What would make a wrong line worse for one recipient than another?

Monday Morning

16 min
What this part is for

Purpose

Ends where the job starts. The candidate owns the workflow from today, and the question is what they stop, what they keep, and who hears it from them rather than from a prospect.

•

It is Monday and the workflow is yours. What do you turn off today, what do you leave running, and who do you tell?

What this question is for, and what to listen for

Purpose

Ownership is read here, in its least glamorous form: making removal requests stick, stopping contact with customers, finding the cause of the refused mail, and telling the people who will be asked about it. A candidate who only lists technical fixes has left the sellers to find out from their customers.

Signals to score

  • Asks what has happened to the 88 people who asked to be removed, and makes their removal hold across every mailbox, domain and the CRM
  • Stops sending from the refused domain, checks the other two, and fixes the shared cause before moving any volume to them
  • Excludes contacts at customer and open-opportunity accounts by account, not by email address
  • Tells the owners of the affected customer and open-opportunity accounts this week, with the names, rather than leaving them to hear it from a customer
  • Stops the same people being enrolled again, and says how long an account rests after it has been emailed
  • Asks the general counsel which rules apply to cold email in each country the contacts are in, rather than deciding it themselves
  • Keeps the parts that are producing qualified meetings running while the rest is fixed
  • Writes down what the workflow does and who to call when it breaks, since the contractor left nothing
  • Adds a check that would have caught each of this month's problems, and says who is told when it fires

Follow-up questions

  • Which of these do you do before lunch?
  • What has happened to the 88 people who asked to be removed?
  • Why not move the refused domain's volume to the other two?
  • What do you tell the account manager of a customer who got this email?
  • What would have told you about the refused domain before week three?

Closing

That is what I had. What I will write up is your Monday list, and I will be honest with you that some of it is ours to fix whether or not you join. We will come back to you within three working days either way.

What this line is for

Purpose

Admits the workflow's problems are the company's, which is true, and gives a dated next step.

Their questions for you, and what happens next. The standard closing

GTM Engineer interviews — common questions

Who is this GTM Engineer interview plan for?
It is written for the interviewer, not the candidate: the hiring manager, engineer or panel member running the Work sample — the outbound workflow's first month round for a GTM Engineer role. It gives you a 60 min script to follow in the conversation — 4 questions with what each one is for and the signals to score against — so you are not writing the round from scratch the night before.
What does the Work sample — the outbound workflow's first month round assess?
This round is focused on: Whether the candidate reads an automated outbound workflow by what it produced rather than what it reports: taking a reply rate apart, finding out how much pipeline would have happened anyway, tracing a wrong personalised line to the data behind it, and deciding what to stop on the first morning they own it. It works through Reading the Pack, Would It Have Happened Anyway, The Line the Model Wrote and Monday Morning, scoring against 36 observable signals, with follow-up prompts on all 4 questions for going deeper where an answer is thin.
How is the 60 min split up?
60 min on 4 questions. The questions take in Reading the Pack (16 min), Would It Have Happened Anyway (14 min), The Line the Model Wrote (14 min) and Monday Morning (16 min). The timings are there so the round stays on schedule and every candidate gets the same shape of interview — which is what makes two candidates comparable afterwards.
What other rounds should I run for a GTM Engineer?

A single round does not cover a whole role. The other rounds in this library for a GTM Engineer:

Hiring for this role?

Open this plan in Hirezen and make it a position in one click.

  • Every interviewer runs the same script and marks the same signals.
  • AI drafts the write-ups, and the debrief puts every read side by side.
  • No ATS to set up first, and no bot in the call.
Start free with this plan

Free for your first open role.