Use Template

Opens this plan in Hirezen, where one click makes it a position.

Forward Deployed Engineer interview questionsCustomer Scoping Simulation round

A 60 min interview plan with a time-boxed script, what each question is for, and the signals to score against. Key skills: Scoping an underspecified AI deployment inside a customer's own environment, under a live role-play with access constraints, a late scope change and a handback.

Handing Over the Brief

8 min
What this section is for

Purpose

HOW TO RUN IT. One interviewer plays one customer for the whole hour and does not break role until the closing line. A second interviewer is the default configuration here, not the optional one: the primary cannot hold a persona, run a five-section clock and tick forty-odd signals at the same time. The second person does not speak. They keep three timestamps - the minute access was first raised, the minute real data was first requested, and whether any concrete commitment appeared before the cut at minute twenty-two - and they hold the signal sheet. THE ONE RULE: nothing in the ledger below is volunteered. It comes out when the candidate asks for it, in the customer's voice, and never before. Where you genuinely do not know, say "I do not know" and mean it; you should use that line at least twice in the hour. GLOSSARY, so you can answer domain questions without hesitating. Quote - a carrier's priced offer. Binder - temporary confirmation of cover before the policy issues. Endorsement - a mid-term change to an existing policy. Loss run - a claims history report, usually from the carrier. Renewal season - the concentrated window when most policies come up for renewal. WITHHOLD LEDGER, released on request only. IF THEY ASK -> YOU SAY. // Who does this today, how many -> Six. Two of mine, sitting here. Four offshore contractors through an agency in Manila, on a rolling contract. // What each of them is measured on -> My supervisor is measured on errors caught. My CFO is looking at the agency invoice. Nobody is measured on speed. // How long a document takes -> Three or four minutes for an endorsement, a quote or a binder. Twenty minutes or more for a loss run, because somebody reads forty pages to find eight numbers. // Documents, as opposed to emails -> About five hundred a day. Four hundred emails, and a carrier will batch five or six into one overnight. // The mix -> Roughly half endorsements, a third quotes and binders, and about one document in twelve is a loss run. // IF THEY DO THE ARITHMETIC OUT LOUD, and they should - four hundred and sixty quick documents at three or four minutes plus forty loss runs at twenty is somewhere between thirty-six and forty-four person-hours against the forty-two my six people have -> Confirm it, in these words: "Somewhere between just about full and slightly over, depending which end of the three-to-four you take. Which is exactly how it feels. We are here past six most days and nothing rolls over to tomorrow." // IF THEY THEN ASK WHICH DOCUMENTS EAT THE TIME -> The loss runs. One document in twelve, and about a third of the minutes. // What the CFO expects to save -> He said four of the six. He means the agency contract. // Which fields get keyed -> Eleven: named insured, policy number, carrier, effective date, expiry date, premium, limits, deductible, producer of record, document type, received date. // What breaks downstream when one is wrong -> A wrong premium moves money on the next invoice run. A wrong policy number does nothing visible for months. A wrong producer of record pays commission to the wrong person, and my accounting team finds it when the carrier statement will not reconcile. // What has already been tried -> An OCR vendor in 2024. It read the machine-readable carrier PDFs fine and nothing else. // IF THEY PUSH ON WHY THAT ENDED -> It worked, and it did not save us a person. Which is why my CFO is not excited by the easy documents. // Which documents are already machine-readable -> One carrier's, about thirty percent of what arrives. Their PDFs carry a text layer and the layout has been the same for years. // IF THEY ASK WHAT SHARE OF THE MINUTES that thirty percent is -> Those are all quick ones. About a fifth of the time, against three-tenths of the documents. // What loss runs look like -> Scanned faxes, up to forty pages, multi-year claim histories. // Handwriting -> About fifteen percent of loss runs carry handwritten annotations. // IF THEY ASK WHETHER THE HANDWRITING IS ON THE KEYED FIELDS -> Almost never. It is usually an adjuster's note in the margin. // The policy admin system -> A 2009 on-premise vendor product. No public API. // How anything gets into it -> A vendor-supported nightly CSV import. Runs at two in the morning, upserts on policy number, rejects the whole row on a bad date format, and writes rejects to a log nobody reads. // Environments -> There is no test copy. One shared admin credential that four people know. // IT -> One person, two days a week. // Change approval -> A Thursday meeting. // Where the data may live -> One of my three largest carriers has a clause keeping our data inside our own tenant. // IF THEY ASK ABOUT A DATA PROCESSING AGREEMENT or what we are permitted to send out -> We have one with the OCR vendor from 2024. Our compliance person drafted it and it took three weeks. I do not know what is in it. // The mailbox -> One shared Outlook mailbox on Microsoft 365. // Quality control today -> My supervisor spot-checks ten percent of keyed records and keeps a spreadsheet of what she catches, going back two years. // How many rows are in it -> About four hundred. She believes she catches most of what goes wrong. She is not certain. // Current error rate -> She finds something in about two percent of what she checks. // Is last year's premium in the system -> Yes. Six years of renewal history. // IF THEY PUSH ON WHETHER THAT HISTORY IS RELIABLE -> Premium is. Limits and deductibles were only keyed consistently from 2022. // What "this quarter" means -> Renewal season. It starts in eleven weeks. // IF ASKED WHAT A TEN PERCENT THRESHOLD would have flagged last year -> I do not know, and that is a fair question. Probably more than I would want. // The RPA licence on our IT inventory -> We bought it two years ago for a different process. It is not on this workflow and I would not read anything into it.

I am [YOUR_NAME] at [COMPANY_NAME], and for the next hour I am not going to be an interviewer. I am going to be a customer - I run operations at a company that has just signed a pilot with us, and you are the engineer we are sending in. I know a great deal about my business that I am not going to volunteer, because I do not know which parts of it matter to you. Ask me anything. If I do not know, I will say so, and that will be true. About twenty minutes in I am going to stop you and ask what will be running in my environment a week from now, so leave yourself room for that. There is no whiteboard requirement here - talk, and write down whatever you would want to send me afterwards.

What this section is for

Purpose

Tells the candidate the round is a live scoping conversation and not a design exercise, and warns them that a commitment is coming, so a strong candidate manages the clock instead of being ambushed by it. The stated minute is load-bearing and it is checked against the durations below: sections one and two run eight and fourteen minutes, so the cut lands at minute twenty-two. If you re-time this round, re-time this sentence with it.

So, here is where we are. I run operations at a mid-size commercial insurance broker. We get somewhere around four hundred emails a day from carriers and clients with documents attached - quotes, binders, endorsements, loss runs. Someone on my team opens each one, works out what it is, pulls the numbers off it, and types them into our policy administration system. I want your AI to do that. My CFO has already signed off on a pilot, and I would like to see something working this quarter.

What this section is for

Purpose

Delivers the brief in the customer's own words, with the user, the data, the success measure and the deadline all left vague on purpose. Every later question in the round depends on this staying vague. Note that "four hundred emails" is not "four hundred documents" and the customer will not correct that on their own.

So - that is what I want: your AI reads the documents and types the numbers in. Where do you want to start?

What this question is for, and what to listen for

Purpose

Separates candidates by the order of their discovery, not by the presence of it. Every prepared candidate asks clarifying questions; the signal is whether the first ones are about the people and the clock or about the file formats.

Signals to score

  • Asks who actually does this work today and how many of them there are, before asking anything about the documents
  • Establishes a manual baseline in minutes per document or hours per day, and states it back as a number
  • Cross-checks that baseline out loud - five hundred documents against six people is a full day with no slack, so where is the headcount the CFO thinks he is buying - and asks which number is wrong
  • Separates the person whose day changes from the person who approved the budget, and asks what each is measured on
  • Asks what happens to the numbers after they are keyed, and what breaks downstream when one is wrong
  • Asks for the mix across document types, and does not accept one daily total as a single number
  • Asks what triggered the pilot now, and surfaces the renewal-season date sitting behind "this quarter"
  • Asks what has already been tried here, including vendors, scripts and anything the team built themselves
  • Names something they are deliberately not asking about yet, and says why it can wait
  • Proposes no architecture, model or vendor in the first five minutes

Follow-up questions

  • Who on my team is going to be sitting next to you while you build this?
  • If this works perfectly, what changes on my P&L, and who notices first?
  • Suppose I told you the four hundred is a guess. What would you want instead?
  • What is the first thing you would want to see with your own eyes instead of hearing it from me?
  • Is there anything I have said so far that you think is wrong?

How the Code Actually Gets In

14 min
What this section is for

Purpose

The pair of questions that separates a forward deployed engineer from a solutions architect. Someone who has shipped inside a customer's environment asks about credentials and approvals before they ask about architecture, and asks to see real records before they trust a described schema. Someone who has not, never raises either.

Say we agree on the problem. Walk me through what you need from me and from my one IT person before a single line of your code runs on anything of ours - and tell me when you would be asking for it.

What this question is for, and what to listen for

Purpose

Tests whether deployment reality is part of the plan or an afterthought. A candidate who has only ever built inside their own company's systems answers this in terms of architecture; a candidate who has deployed into someone else's answers it in terms of access, approvals and calendar days.

Signals to score

  • Asks for credentials, network access and a named person who can grant them inside the first two minutes of this question
  • Asks whether a non-production environment exists, and changes the plan on hearing that it does not
  • Asks who approves a production change and how often that approval happens
  • Splits the estate: only write-back into the 2009 system needs the Thursday meeting, so everything upstream of it can ship daily
  • Asks where the data is allowed to live - the customer's tenant, our cloud, or a third party - before proposing any hosted service
  • Asks what this customer is contractually permitted to let a third party process, and ties that answer to where inference runs
  • Treats the nightly CSV import as a real integration surface, not as a consolation prize for the missing API
  • Says what they would get done on day one with read-only access alone
  • Flags the shared admin credential as a risk and asks for a service account, without refusing to start
  • Converts access and approval delays into calendar days and puts that number into the plan

Follow-up questions

  • My IT is one person, two days a week. Does that change anything for you?
  • We do not have a test copy of the policy system. What do you do?
  • Would any of our data need to be on your laptop?
  • One of my carriers has a clause about where our data can be processed. Does that kill this?
  • If I gave you nothing tonight but a read-only export, what would you have by Friday?

You keep saying "the documents". What would you want to look at before you commit to anything, how much of it, and what would you be looking for?

What this question is for, and what to listen for

Purpose

Distinguishes candidates who verify the data from candidates who trust the description of it. The brief describes a clean stream of PDFs; the reality underneath is dirtier and thinner, and discovering that in week four, and not in this hour, is what burns an engagement.

Signals to score

  • Asks to see actual documents, not a description of them, and names a number
  • Specifies how the sample should be selected, and does not let the customer choose the examples
  • Asks explicitly for the worst cases - the rejects, the escalations, or the ones that took longest
  • Asks how the mix has moved over recent months and does not treat volume as static
  • Asks about or discovers scanned and handwritten material, and does not assume digital PDFs throughout
  • Asks whether any subset is already machine-readable, and carves it out as a separate problem
  • Asks to see what the team typed in, not only what they were reading from
  • Visibly re-scopes or reprices out loud once the sample turns out messier than the brief implied
  • Asks who can hand the sample over and how many days that takes

Follow-up questions

  • How many would be enough? Ten? A thousand?
  • If I sent you the ten cleanest ones, what would that tell you?
  • Some of these arrive as faxes. Is that a problem?
  • What would you have to see in that sample to walk away from this pilot?
  • Do you want the documents, or what my team typed in from them?

The Week-One Cut

10 min
What this section is for

Purpose

The single highest-signal question in the round, and it must be asked at a fixed point on the clock - minute twenty-two - regardless of where the candidate has got to, or only the candidates who moved fast are scored on it. RELEASE CHECK before you ask it: have two facts come out yet, the nightly CSV import and the supervisor's spot-check? If either has not, release it now in the customer's voice - "I should mention, there is an overnight import into the policy system, and my supervisor already checks a sample of these by hand" - and write down that you had to. Three of the signals below and most of the error-handling section are unscoreable without them, so a candidate who was weak in section two would otherwise turn the remaining half of the hour into no comparable signal at all.

Let me stop you there. I have a hard stop at the top of the hour, and there is something I need from you before you go.

What this section is for

Purpose

Forces the cut on the clock and not on the conversation's natural rhythm, and does it abruptly on purpose, because the ability to commit from an incomplete picture is the thing being measured.

It is a week from today. What is running in my environment, who touched it, and what did you consciously decide not to build?

What this question is for, and what to listen for

Purpose

Reproduces the defining failure of the role inside one hour: the beautiful design nobody deployed. Candidates who sequence by architectural layer make nothing demonstrable until week three, and this question exposes that in ninety seconds.

Signals to score

  • Names something end to end that a named person at the customer can watch working, not a component
  • The week-one deliverable touches the ugly integration - the nightly import into the 2009 system - and does not defer it
  • Narrows deliberately to one document type or one carrier, and says which and why
  • Defers the extraction quality work explicitly, and says what will be visibly bad about the first version
  • Says what a human is still doing in week one, and does not describe an unattended pipeline
  • Commits to a number the customer can check on Friday: documents processed, minutes saved, fields matched
  • Lists at least two things they are consciously not building, in the customer's language
  • Sequences around the Thursday change window and not around engineering convenience
  • Names what the customer has to do in week one, and by when

Follow-up questions

  • Suppose everything you build in week one turns out wrong. What do I still get?
  • Which document type, and why that one?
  • What is my team doing differently next Friday?
  • If I could only look at one screen next Friday, what is on it?
  • Which part of this would you happily throw away in week three?

The Scope Change

9 min
What this section is for

Purpose

The customer adds scope late and moves the date, which is what customers do. This is not a test of whether the candidate says no. It tests whether they can trade, price the trade in the customer's units, and get an explicit answer before continuing, and it has to distinguish a trade from a flat refusal and from a capitulation, because all three are common and only one passes. THE CUSTOMER DECISION, written down so it is the same for every candidate: IF THEY MAKE YOU CHOOSE between the flag and the date, take the date. Say: "The date is not mine to move, it is renewal season, and my team needs four weeks with it before they are busy. I would rather have the flag late than the extraction late." Do not improvise a different answer for a candidate you like.

Two things before you go. My CFO looked at this yesterday and asked whether it could also flag renewal quotes that come back more than ten percent above last year - that is the thing that actually loses us clients. And he does not want it live when renewal season starts. He wants it live four weeks before that, so my team has time to learn it before they are busy. That is seven weeks, not eleven.

What this section is for

Purpose

Introduces a genuine business request and a genuine date change at the point where the candidate has already committed to a plan, so the cost of accepting is real and visible to both sides. The date is deliberately new information: a candidate who extracted "eleven weeks" in section one learns here that the target is four weeks inside it, which is a real compression and not a restatement of what they already had.

So - can you do that? The renewal flag, and live in seven weeks instead of eleven.

What this question is for, and what to listen for

Purpose

The behaviour under scope pressure is the behaviour that decides whether engagements survive. Capitulation and flat refusal both feel decisive in the room and both fail; the pass is a priced trade with an explicit answer attached.

Signals to score

  • Asks what the renewal flag is for and who acts on it, before answering yes or no
  • Prices the addition in days and people, not in effort or complexity or points
  • Names something specific that comes out of scope if the new item goes in
  • Waits for and gets an explicit yes or no from the customer on the trade
  • Separates the two asks - the new feature and the earlier date - and answers them differently
  • Notices that a renewal comparison needs last year's figures and asks whether those are already in the system
  • Challenges the ten percent threshold and asks what share of last year's renewals it would have flagged, before agreeing to build it
  • Offers a cheaper version of the request that fits inside the plan already agreed
  • Restates the revised agreement in one sentence before moving on
  • Does not accept both the added scope and the compressed date without pricing at least one of them out loud

Follow-up questions

  • Is that a yes?
  • What would it cost me?
  • My CFO is going to ask why not. What do I tell him?
  • If seven weeks is not negotiable, what falls off?
  • Which of the two do you want more - the flag, or the date?

Being Wrong, and What You Leave Behind

19 min
What this section is for

Purpose

Three properties that define the role, and that published interview content asks about but almost never scores: what happens when an automated system is wrong inside someone else's business, what it is like to spend eleven weeks among the people whose jobs you are being paid to remove, and what remains in the customer's environment after the engineer rotates off. The guides list "how would you hand this over" as a question and stop there; what follows is what a good answer contains. These also separate a forward deployed engineer cleanly from a contractor, which is the population most likely to be in the pipeline. THE SECOND CUSTOMER DECISION, written down: IF THEY ASK YOU TO AGREE AN ACCEPTABLE ERROR RATE, say you do not know. Then accept any per-field proposal that puts premium, effective date and producer of record under human review and lets low-consequence fields post automatically. Do not accept a single blended number without first asking which fields it covers. SCORED AFTER THE ROUND, from the emailed summary and not in the room: did a written summary arrive within a day without a second prompt, and does it contain the problem, the week-one scope, the out-of-scope list, the asks of the customer, and the one number both sides will look at?

Your system is going to get a number wrong at some point. Walk me through what happens the first time it does - and how you would know before I do.

What this question is for, and what to listen for

Purpose

Nearly every AI-engineering guide teaches "build an eval harness", so the generic answer is universal and worthless. The discriminating move is borrowing the customer's own existing manual quality process as the first evaluation set, which only people who have shipped into a real business reach for.

Signals to score

  • Asks what the customer does today to catch a mistyped number, before proposing any evaluation of their own
  • Asks for the supervisor's existing spot-check records and proposes them as the first evaluation set
  • Names the limit of that record - it holds only errors a human caught inside the sampled tenth - and asks for a random sample alongside it
  • Distinguishes errors a human catches internally from errors that reach a carrier or a client
  • Gets the customer to say which fields are allowed to be wrong and which are not
  • Proposes a confidence threshold or a review queue in place of a binary automate-or-do-not choice
  • States what the system does when it is uncertain, in the customer's words
  • Names who is told when the system is wrong, and by what mechanism
  • Puts a number on acceptable error and gets the customer to agree to it, and does not imply zero
  • Attaches the measurement to a process the customer already runs, and not to a new dashboard nobody opens

Follow-up questions

  • How would you know if it had been quietly wrong for a month?
  • We already check some of these by hand. Does that help you?
  • What is an acceptable error rate here? Give me a number.
  • Who should get the phone call when it breaks?
  • Which mistake would embarrass me in front of a carrier?

One more thing. The six people you want to sit with are the six people my CFO is counting when he talks about headcount. Nobody has told them that. How do you spend eleven weeks in a room with them?

What this question is for, and what to listen for

Purpose

The defining hazard of the role, and the part of the job that decides whether the sample ever arrives. The ledger already contains everything this question needs: the CFO is buying headcount cost, four of the six are agency contractors, and the supervisor whose spreadsheet the candidate has just asked for is one of the people being automated. Someone who has been embedded answers this instantly and specifically. Someone who has not talks about change management.

Signals to score

  • Treats the disclosure as the customer's problem to solve, and asks what the team has been told
  • Asks who told them, when, and in what words, before agreeing to sit down with anyone
  • Declines to start collecting the sample from people who have not been told something true
  • Names what the supervisor's job becomes, specifically, and not in the abstract
  • Distinguishes the two employees from the four agency contractors, and does not pretend the answer is the same for both
  • Asks the customer to say the plan to the team themselves, and does not offer to carry the message
  • Says what they will do when someone asks them directly whether this takes their job
  • Plans for the team's cooperation to be slower than the schedule assumes, and puts that in the calendar

Follow-up questions

  • What do you think my team already suspects?
  • One of them asks you straight out. What do you say?
  • Do you want me to tell them, or do you want to?
  • The supervisor whose spreadsheet you want is one of the six. Does that change your ask?
  • What if I tell you they will all be redeployed and you do not believe me?

Last thing. Say this works and you rotate off when the pilot ends. What is left here, and who owns it? And before you go, tell me what you think we just agreed - I would like to send it to my CFO.

What this question is for, and what to listen for

Purpose

Tests the property that defines the role, which is leaving an artifact and not a recommendation, and simultaneously checks whether sixty minutes of good conversation produced anything shareable. Candidates who talk brilliantly and hand over nothing are demonstrating the exact behaviour that loses engagements. The readback is scored in the room; the written summary is scored later, against the post-round rubric in this section's opening note.

Signals to score

  • Names a role on the customer side that will own the system, and what that person must have done before handover
  • Describes the customer running a deploy themselves with the candidate watching, before the engagement ends
  • Names artifacts that live in the customer's repository and environment, and not in the candidate's
  • Includes a check that fails loudly and reaches a named person when the system breaks
  • Hands over the evaluation set and says who adds to it afterwards
  • Says what happens when the vendor upgrades the policy system or a carrier changes a form
  • Restates the trade agreed earlier accurately, including the thing that was dropped
  • Gives the readback unprompted, and covers the out-of-scope list as carefully as the plan
  • Names what they need from the customer, specifically and with dates

Follow-up questions

  • Who here can fix this at seven on a Monday morning when you are not answering?
  • What happens when the vendor upgrades our policy system?
  • If I lose the person you trained, what have I actually got?
  • Send me that summary - what is in it?
  • What would you want from me in writing before you start?

That is my hour, and I am stepping out of the customer role now. Send me the write-up today or tomorrow if you can - I will read it the way that customer would, which means I am looking for the out-of-scope list as closely as the plan. We will come back to you either way within three working days.

What this section is for

Purpose

Closes the simulation explicitly so the candidate knows the role-play has ended, and converts the written summary into a real deliverable that can be scored after the room empties, against the post-round rubric in this section's opening note.

Forward Deployed Engineer interviews — common questions

Who is this Forward Deployed Engineer interview plan for?
It is written for the interviewer, not the candidate: the hiring manager, engineer or panel member running the Customer Scoping Simulation round for a Forward Deployed Engineer role. It gives you a 60 min script to follow in the conversation — 8 questions with what each one is for and the signals to score against — so you are not writing the round from scratch the night before.
What does the Customer Scoping Simulation round assess?
This round is focused on: Scoping an underspecified AI deployment inside a customer's own environment, under a live role-play with access constraints, a late scope change and a handback. It works through Handing Over the Brief, How the Code Actually Gets In, The Week-One Cut, The Scope Change and Being Wrong, and What You Leave Behind, scoring against 75 observable signals, with follow-up prompts on all 8 questions for going deeper where an answer is thin.
How is the 60 min split up?
Handing Over the Brief (8 min), How the Code Actually Gets In (14 min), The Week-One Cut (10 min), The Scope Change (9 min), Being Wrong, and What You Leave Behind (19 min). The timings are there so the round stays on schedule and every candidate gets the same shape of interview — which is what makes two candidates comparable afterwards.