Why we interview a business with a voice agent, not a form
Forms are a guess made in advance. When you are trying to understand a business, a voice agent that can follow the tangent beats a questionnaire that cannot.
Why we interview a business with a voice agent, not a form
Every engagement starts the same way. You need to understand the business before you can help it, so you send a discovery questionnaire. Twenty fields, a few text boxes, a polite note asking them to be as detailed as they can.
Then it comes back, and it tells you almost nothing you can act on.
Someone writes "the handover process is a nightmare" in a text box and moves on. That sentence is the single most important thing in the whole form, and the form has no way to do the one thing any decent interviewer would do next: lean in and ask "what specifically breaks?"
That is the problem with a form. It is a guess made in advance. You sit down before you have met anyone, try to predict every question worth asking, and freeze the list. Whatever you failed to anticipate, you never learn. The person on the other end does the minimum, skips the fields that feel like effort, and the richest material, the aside, the "oh, and the thing that really drives us mad", never gets written down because no box asked for it.
So for our AI-First discovery work, we stopped sending forms. We send a voice agent instead.
What we built
Before we build anything for a client, we need to understand two things: where the founder wants the business to go, and how the work actually gets done by the people doing it. Those are two different interviews with two different people, and neither is well served by a questionnaire.
So each person gets a link. They click it, and a voice agent talks to them for fifteen or twenty minutes. It asks about their work, listens to the answer, and follows up on the interesting bits. It is self-serve and async, so nobody has to schedule a call with us.
When the call ends, the conversation is transcribed, the key points are pulled into structured fields automatically, and everything lands in one place. Out of all those conversations we build an Impact Map: a picture of where this specific business could actually use automation, drawn from the people who do the work rather than from us guessing on their behalf.
The output is the part that surprises people. You get the depth of a real conversation and the structure of a form. You do not have to choose between them.
Why it is easier for them, too
There is a quieter reason this works, and it has nothing to do with data quality. It is simply easier for the person on the other end.
A form is homework. A scheduled call is arguably worse: now everyone has to find a slot, be at a desk, and be "on" at 2pm on a Thursday whether that suits them or not. A lot of the meetings booked for this sort of thing could have been an email, and everyone sitting in them knows it.
The voice interview asks nothing of the calendar. There is no meeting to arrange, no room, no time that three busy people have to line up. Each person just gets a link and does it whenever actually suits them. On their phone while walking the dog. With a coffee at five in the morning. Quietly one evening once the house has gone to sleep. Whichever fits their day.
And talking is less effort than writing. Most people can say far more than they will ever type into a text box, and they tend to say it more honestly, because it feels like a conversation rather than a test. You get more out of them, and it costs them less to give it.
The payoff comes back to them, too. Because every minute of it feeds straight into what gets built for their business, it is not another hour that vanishes into a calendar. It is the rare bit of "admin" that actually returns something to the person who did it.
How it works
The interviewer is an ElevenLabs voice agent running on a small, fast model (we use Claude Haiku, which we found listens better and marches through a question bank less robotically than the alternatives we tested). Each person gets a durable web link that embeds the agent directly in a branded page, so there is no app to install and no third-party login. When the call ends, ElevenLabs runs its own analysis, extracts a set of predefined fields from the transcript, and fires a webhook. A Make scenario catches it and writes the summary, the full transcript, and every structured field into the matching row of an Airtable base.
The important design decision is that there is not one interviewer. There is one per role: founder, sales, operations, finance, project delivery, marketing. They share a spine, but they are not the same agent, and the reason why is the whole game.
The founder is the source. Everyone else is the subject-matter expert on their own work.
The founder interview captures vision and direction. The operator interviews capture reality: how the job is really done, where it hurts, what they would fix if they could. So we feed each operator agent a short briefing about the company, the founder's vocabulary, the direction of travel, so it walks in sounding like someone who knows the business. But we deliberately withhold one thing: the founder's opinion of that person's work.
If the agent turns up carrying the founder's view of how the operations lead "should" be working, it stops discovering and starts confirming. It leads questions to prove the founder right, frames the person as a deviation from a baseline, and, worst of all, the person can feel it and gets defensive. The whole point of interviewing the people who do the work is that the founder does not have full visibility of the day-to-day. That gap is exactly what you are trying to fill. Poison it with the founder's assumptions and you have built an expensive form.
The gotcha that taught us the most
Injecting context into an agent is easy to get subtly wrong. On one early build we gave the finance interviewer a context block that mentioned the founder's top priority, which happened to sit around proposals and costings. We told the agent to open by anchoring on "something concrete you already know about their role."
The agent did exactly that. It opened by telling the finance person they handled the sales proposals process. That was not their job at all.
The lesson generalised into a rule we now apply to every agent that gets injected context: do not ask the model to choose what to anchor on. Give it the exact opening sentence, written per role, and tell it to use that verbatim. The moment you let it pick the "most concrete" thing out of a block of context, it will occasionally pick the wrong thing and say it with total confidence. Precision in, precision out.
Making it sound like a person, not a phone tree
A voice agent that reflects every single answer back at you ("so what I'm hearing is...") is exhausting inside two minutes. Getting the conversational texture right took a specific set of changes, applied together:
- An acknowledgement mix: roughly half bare acknowledgements, a third short synthesis, and only occasionally a full playback. Telling the model "don't paraphrase every turn" did nothing. Giving it percentages changed the behaviour.
- Probe before pivoting. If someone trails off or gives a thin answer, ask one short "anything else there?" before moving on, rather than stacking a paraphrase and a fresh question on top.
- Quiet progress signposting ("about a third of the way in"), because the model has no real sense of elapsed time, and neither does the person if you never tell them.
- A verbal escape hatch in the opening line: if I move on too quickly and you had more to say, just say "one more thing". People use it.
- A structured close that says what happens to the interview, what comes next, and when they will hear from us, instead of hanging up abruptly.
None of those are independently optional. They reinforce each other. Probe-before-pivot only feels natural when the agent is not paraphrasing constantly, and the close only lands if the pacing stayed calm the whole way through.
The interview that improves itself
Here is the part a form can never do: at the end, the agent asks how it did.
Not a satisfaction score. A genuine question about the experience, what felt natural, where it moved on too fast, where it missed something. And because the person has just spent twenty minutes talking to it, they have real opinions.
That feedback turned out to be one of the most valuable outputs of the whole process, because it tells you exactly where the agent's behaviour breaks. The five changes in the section above did not come from us theorising in a room. They came from reading back what people told us about their own interviews. Someone said the ending felt abrupt and left them unsure what happens now, so we built the structured close. Someone noticed it paraphrased everything and it grated, so we built the acknowledgement mix. Five or six changes so far, each one traceable to a real interview that told us what was off.
A form is frozen the day you write it. Every person who fills it in gets the same twenty fields, and the thousandth submission is no better handled than the first. The agent goes the other way. Every interview is a chance to find where it fell short and tune it, so the next one is a little sharper. The version running in month three is meaningfully better than the one we shipped in month one, and we did not have to guess at the improvements. The people who used it told us.
That is the quiet advantage. You are not just capturing better information than a form would. You are running a tool that gets better at capturing it every single time.
What you can take from this
The honest version of this is not "voice beats forms". It is sharper than that.
A form is the right tool for a transaction. A conversation is the right tool for understanding.
When someone already knows exactly what they want, a booking, a date, a quantity, a form is faster and a conversation is just friction. We build plenty of forms, and we would never put a voice agent in front of someone trying to hire a vehicle for Tuesday. The structured, known, transactional stuff is precisely what forms are for.
But the moment your real goal is to understand a person's situation rather than capture a known input, the form becomes the bottleneck. Discovery is the clearest case. So is onboarding, feedback, and any time the most valuable answer is one you did not know to ask for.
A few principles carry across, whatever you build:
- The value is in the second question. Whatever method you use, if it cannot follow up, it will miss the thing that mattered. Design for the tangent.
- Give structure to the tail end, not the front. A form forces structure before the conversation, which caps how much you can learn. Let the conversation run, then extract the structure from the transcript afterwards. You get both.
- Mind whose assumptions you are carrying. A tool that walks in already "knowing" the answer will find the answer it expected. Whether that is a leading form field or an over-briefed agent, the failure mode is identical.
- Build the feedback loop in from the start. A form is frozen the moment you ship it. Anything that can ask how it did, and be tuned from the answer, improves with use instead of standing still. Design the thing to get better.
Closing thought
The irony is that the technology here is not really the point. The voice agent is just a way to do, at scale and without booking anyone's calendar, the thing a good interviewer has always done: stop talking, listen, and ask the obvious follow-up. The form never could, and it never improved. This does, a little more with every interview. We just stopped pretending the form was good enough.
If you are trying to map where AI could genuinely help your own business, that mapping is the work we do first, and we do it by talking to your team, not by handing them a questionnaire.