Menu
← FIELD NOTES
TAGGEDVOICE 10 POSTS

Posts about VOICE.

2026.09.14 VOICE

You scrubbed the words from the audio and shipped the speaker's age, accent, and health anyway.

Redacting the transcript and anonymizing the speaker feels like the privacy work is done. It is not: the waveform is a biometric giving up age, accent, emotion, and clinical conditions from the signal — and it leaks hardest for the groups least seen in pretraining.

2026.09.11 VOICE

Speech deepfake detection is losing the arms race.

Anti-spoof detectors hit near-perfect scores in the lab and collapse in the open world. They memorize the artifacts of the attacks they were trained on, and every new synthesizer erases them. Detection has to be one layer — paired with liveness checks and out-of-band verification.

2026.08.27 VOICE

Your streaming transcriber is a worse transcriber, and the benchmark you bought it on was offline.

The WER on the ASR model card was measured with the whole utterance in hand. Your real-time voice agent never gets the whole utterance — it runs a different, worse operating point of the same weights, and no latency ledger records the bill.

2026.08.06 VOICE

Voice agents need joint audio-text safety.

Safety alignment trained on text does not fully transfer to voice. Audio-native attacks, and joint audio-text attacks that exploit the seam between modalities, walk through guardrails that hold firmly against the same content in writing.

2026.07.07 VOICE BUILD NOTE

Turn-based voice RAG: the whole loop on one stack.

A mic button on a grounded-answer dialog: Whisper for transcription, the unchanged retrieval path, per-sentence TTS for playback. What broke, what it costs, and why the label says turn-based instead of realtime.

2026.06.29 VOICE

Voice agent evals: scoring a conversation, not a transcript.

A voice agent can pass every transcript-level metric — low word error, correct answers — and still be unbearable to talk to. The caller experiences timing, turn-taking, and recovery, none of which a transcript records. How to score the conversation.

2026.05.15 VOICE

Cascaded or end-to-end: a 2026 voice-architecture trade study.

A 2026 voice agent forks at the first design decision — STT→LLM→TTS cascade, or a single speech-to-speech model. End-to-end wins on latency and naturalness; the cascade wins on everything you debug, audit, and control. Here is the trade study.

2026.05.02 VOICE

Graceful failure for voice agents.

A voice call is real-time and unforgiving — there is no spinner to show, and dead air reads as a broken product. When STT, the LLM, or TTS fails mid-call, the system has to degrade, not drop.

2026.04.07 VOICE

Turn-taking is the hard part of voice agents.

Transcription is largely solved. Knowing when the caller has finished, when to stop for an interruption, and when an 'mm-hm' is not a turn — that is not. Endpointing, barge-in, and backchannels, measured.

2026.02.15 VOICE

A 280ms latency budget, broken down millisecond by millisecond.

Sub-300ms voice agents are a specific engineering problem. Here is every millisecond a packet spends between the user's mouth and the agent's reply — and where you actually claw the time back.

NEW ENGAGEMENT · INTAKE

Tell us about it.

The more specific you are, the more useful our first reply.

SERVICE AREA
↩ ENCRYPTED IN TRANSIT
ASK THE FIELD NOTES BETA