Menu
← FIELD NOTESVOICE 2026.06.29 · 10 min

Voice agent evals: scoring a conversation, not a transcript.

A voice agent can pass every transcript-level metric — low word error, correct answers — and still be unbearable to talk to. The caller experiences timing, turn-taking, and recovery, none of which a transcript records. How to score the conversation.

Two artifacts come out of the same phone call. One is the transcript: a column of speaker labels and the words each said, the thing that lands in the eval harness and gets scored. The other is the recording: the same words, but now carrying every pause, every overlap, every half-second where nobody spoke and the caller wondered if the line had dropped. The eval reads the first. The caller lived the second. A voice agent program that scores only the transcript is optimizing the artifact nobody experiences and ignoring the one everybody does.

That is not a subtle failure. An agent can post a low word error rate, return correct answers on every task in the suite, and clear every gate the team set — and then go to production and get hung up on. It answers correctly two seconds after the caller stops speaking. It talks over people. It leaves a long silence with no signal that it is still listening. None of those defects show up on the page, because the page has had time stripped out of it, and time is where they all live. This essay is about the metrics that survive being written down and the ones that do not, and why a serious voice eval has to score both.

A transcript is a recording with the time deleted

Start with what the transcript actually is, because the word makes it sound more complete than it is. A transcript is a lossy projection of a conversation onto text. It keeps the words and the speaker order. It discards everything continuous: the gap between turns, the overlap when two parties speak at once, the duration of a silence, the rising pitch that signals a question is not finished. Those are not decorations on top of the words — for a phone call they are load-bearing, and the projection throws them all away.

Word error rate and answer correctness are computed on that projection, which is exactly why they are necessary and exactly why they are not sufficient. They are necessary because an agent that mishears the account number or gives the wrong balance has failed, and you want a number that catches it. They are insufficient because they can only ever score the dimension the projection preserved. WER asks whether the right words were recognized. Answer accuracy asks whether the right words were produced. Neither can ask when — and “when” is the axis the caller is most sensitive to.

So the claim of this essay is direct, and it does not need hedging: word error rate and answer accuracy measure a real but narrow slice of a voice agent, and a program built only on them will systematically ship agents that score well and converse badly. The benchmarks in the sections below each measure something the transcript structurally cannot hold. Read together, they are the same claim from the constructive side — here is what to score instead.

THE TRANSCRIPT KEEPS              THE CONVERSATION ALSO HAS
┌──────────────────────┐          ┌────────────────────────────────┐
│ words spoken         │          │ gap before each turn (ms)      │
│ speaker order        │ ──────►  │ overlap / who interrupts whom  │
│ answer correctness   │  time    │ silence duration mid-turn      │
│                      │ deleted  │ backchannels ("mhm") + timing  │
│                      │          │ recovery after a mishearing    │
└──────────────────────┘          └────────────────────────────────┘
  WER + accuracy scores             EVA-X-style experience metrics
  this column                       score this column

The 200-millisecond gap is an empirical fact, not a preference

The first thing the projection deletes is the gap between turns, and that gap has a measured human value. Studies of conversation across ten languages (Stivers et al., PNAS, 2009) found the modal gap between one speaker finishing and the next beginning is around 200 milliseconds, with cross-linguistic variation sitting inside a band of roughly 250 milliseconds. The number is remarkably stable across languages that differ in almost every other respect, which is the tell that it is not a stylistic habit. It is a property of the human turn-taking system — and a listener’s expectation is calibrated to it without the listener ever being aware of the calibration.

That calibration is why latency reads as wrongness rather than as slowness. When an agent’s response gap sits far outside the human band, the caller does not consciously clock the delay in milliseconds. They feel that something is off — that the agent is distracted, or did not understand, or is not really there — and they react to that feeling, often by repeating themselves or starting to talk over the agent, which then degrades the next turn too. A timing defect is not a clean isolated penalty; it propagates into the turns after it.

Be precise about the number, though, because it is routinely mis-stated as a hard pass threshold. The 200 milliseconds is a human modal gap, not a line below which an agent is fine and above which it is broken. Production systems run slower than that and remain perfectly usable; practitioner budget breakdowns from real deployments treat roughly 300 milliseconds as a good target and locate the onset of trouble past about 500, where callers begin to repeat and collide. Full-duplex models such as Moshi (arXiv 2410.00037) report theoretical latency near 160 milliseconds precisely because closing on the human band is the design goal. For the eval the consequence is concrete: response latency is not a single average to glance at. It is a per-turn measurement and a distribution to score — a long tail of slow turns can hide behind a healthy mean, and the tail is what the caller remembers.

Turn-taking: the axis where identical transcripts diverge

The second casualty of the projection is the exchange itself — who takes the floor, when, and whether two parties land on it at once. This is measurable, and recent benchmarks make it concrete rather than impressionistic. Full-Duplex-Bench (arXiv 2503.04721) defines automatic metrics for it: a Takeover Rate for how often the agent seizes the turn, backchannel frequency and the timing distribution of those backchannels — the “mhm” that tells a speaker they are still being heard — response latency, and latency after an interruption. Talking Turns (arXiv 2503.01174) paired a turn-taking evaluation protocol with a user study and found that current spoken-dialogue systems “do not understand when to speak up, can interrupt too aggressively and rarely backchannel.”

Hold that finding against the transcript and the reason turn-taking needs its own axis becomes unavoidable. Picture two agents handling the same caller, who pauses for breath halfway through stating a problem. Agent one waits, lets the caller finish, then responds. Agent two treats the breath-pause as a turn boundary and cuts in. Transcribe both calls and the transcripts are nearly identical — the same words, in the same order, from the same two speakers. Score them on WER and answer accuracy and they tie. But one call felt like a conversation and the other felt like being talked over by a machine, and the only artifact that recorded the difference was the timing the transcript deleted.

That is the structural argument for a separate score. Turn-taking quality is not a refinement of transcript accuracy that you could fold into the same number; it is orthogonal to it, invisible to it, and central to the experience in a way transcript accuracy is not. The mechanics of getting it right — endpointing, barge-in handling, back-channeling — are the subject of the voice-turn-taking essay. The point for the eval is narrower and prior to mechanics: whatever turn-taking behavior you built, you have to measure it directly, with metrics like Full-Duplex-Bench’s, because no transcript score will ever surface it.

Recovery: scoring the path not taken

The third thing the projection hides is failure handling, and it hides it for a specific reason: a transcript records the path the conversation actually took, never the recovery the agent did or did not make. Real calls break in ordinary ways. Speech recognition mishears an entity — a name, a date, an address. The caller interrupts mid-sentence. A silence opens and the agent has to decide whether it is a turn boundary or a beat of thought. Whether the agent re-confirms the misheard entity, yields gracefully to the barge-in, or fills the silence with a holding signal instead of dead air is a large part of whether the call survives contact with the real world.

Consider a concrete case. The caller says “transfer it to my savings,” the recognizer hears “transfer it to my checking,” and the agent now holds a wrong entity. There are two continuations. In the first, the agent reads the action back — “moving it to checking, is that right?” — the caller corrects it, and the call recovers. In the second, the agent silently proceeds on the wrong account. Now look at what the transcript of each shows. The first transcript contains a confirmation turn and a correction. The second contains a clean-looking exchange with no visible error at all — the mistake is in the mismatch between what the caller meant and what the agent did, and a transcript has no column for intent. The worse outcome leaves the cleaner transcript.

That is why recovery cannot be evaluated on a happy-path test set. If every call in the suite is a clean call, the suite never exercises the behavior that most determines real-world quality, and it will rate a brittle agent and a robust one the same. A voice eval has to deliberately seed its test set with mishearings, mid-turn interruptions, and ambiguous pauses, and then score each one on whether the agent recovered — not on whether the transcript reads tidily afterward. The mechanics of recovering well are the voice-graceful-failure essay’s subject; the eval requirement is that the failures have to be in the test set on purpose, because they will be in production whether you put them there or not.

Two scores, kept apart

Putting it together gives the shape of the eval: it has to score the call as a call, and it has to do that without collapsing two different things into one number. The EVA-Bench framework (arXiv 2605.13841, a 2026 preprint) is built around exactly this split. It separates EVA-A — accuracy metrics such as task completion and audio fidelity — from EVA-X — experience metrics including conversation progression and turn-taking timing — and it runs the evaluation as bot-to-bot multi-turn audio conversations rather than scoring isolated transcribed turns.

The separation is the part to copy, and it is worth being explicit about why averaging is the wrong move. An agent can be strong on transcript accuracy and weak on conversational experience, or the reverse — the two axes genuinely come apart. Average them into a single quality score and a fast, fluent, interrupting agent and a slow, polite, accurate one can land on the same composite, while a number that should have told you precisely which half is broken instead tells you nothing actionable. Two scores reported side by side keep the diagnosis intact: you can see at a glance whether the next sprint goes to the recognizer and the answer logic or to the timing and turn-taking, and a regression in one cannot be masked by a gain in the other.

Two cautions sit on top of that structure. First, many benchmarks marketed as “voice” — VoiceBench (arXiv 2410.17196) among them — still largely score knowledge and instruction-following on synthesized speech. They are useful for what they measure, and what they measure is not conversation; check what a benchmark actually scores before trusting it to cover the experience axis. Second, wherever the eval uses a model to judge conversational quality, that judge is itself a model carrying biases and drift, and the discipline the LLM-as-judge essay lays out — calibration, spot-checking against human ratings — applies to the voice judge as much as to any other.

A voice eval that scores what the caller hears

The whole argument reduces to a short list of properties a voice agent eval has to have. Each item exists because a transcript-only program will miss the thing it names.

  • Latency is measured per turn and scored as a distribution, against a roughly 300 ms target — not collapsed to a single average that hides a slow tail.
  • Turn-taking is its own scored axis: takeover rate, interruption handling, backchannel frequency and timing — the metrics a transcript cannot expose.
  • The test set is seeded with mishearings, interruptions, and ambiguous pauses, deliberately, because production will contain them and a happy-path suite never will.
  • Recovery from each failure is scored, on whether the agent re-confirmed, yielded, or held the line gracefully — not on whether the transcript reads cleanly afterward.
  • Transcript-accuracy and conversation-experience are reported as two separate scores, never averaged, so a regression in one cannot hide behind a gain in the other.
  • Evaluation runs on multi-turn audio conversations, not isolated transcribed turns, because timing and turn-taking only exist across a continuous exchange.

Reading list

  • Stivers et al. — universals and cultural variation in turn-taking; the ~200 ms human response gap: pnas.org
  • Full-Duplex-Bench — automatic metrics for turn-taking: takeover rate, backchannel timing, latency after interruption: arXiv 2503.04721
  • Talking Turns — a turn-taking evaluation protocol; current systems interrupt too aggressively, rarely backchannel: arXiv 2503.01174
  • Moshi — a full-duplex speech-text model and its ~160 ms latency target: arXiv 2410.00037
  • EVA-Bench — an end-to-end voice-agent eval that separates accuracy from experience: arXiv 2605.13841
  • VoiceBench — a voice-assistant benchmark; useful, and an example of transcript-level scoring: arXiv 2410.17196

Score the recording, not the page — because the caller never reads the page, and the page never recorded the half-second that made them hang up.

NEW ENGAGEMENT · INTAKE

Tell us about it.

The more specific you are, the more useful our first reply.

SERVICE AREA
↩ ENCRYPTED IN TRANSIT
ASK THE FIELD NOTES BETA