Dead Air Is Never One Cause

When a callbot stops talking and never picks back up, the symptom looks identical every time but the cause rarely is. Hamming's analysis breaks the same silence into distinct root causes: an RTP stream that dropped while the application layer still thinks the call is live, voice-activity-detection (VAD) endpointing that waits too long after the caller stops speaking, speech recognition (STT) that keeps emitting partial transcripts without ever finalizing, a tool call stuck mid-lookup, or TTS generation and playback running late. Fixing one layer leaves the silence from every other layer untouched.

That is why Hamming recommends logging every silence as a turn-level event and joining component timestamps — telephony, VAD, STT finalization, LLM first token, tool execution, TTS first audio, and playback — so each gap can be attributed to the layer that actually caused it. Lump every silence into "slow response" and the team that owns the real cause never gets paged.

One Threshold Can't Tell Them Apart

FutureAGI's dead-air check uses RMS energy with defaults of a 20% total-silence budget per call and a 3,000ms maximum single gap. What matters more than the numbers themselves is that they are an evaluation tool's defaults, not a law of physics: a short gap could be the model reasoning, or it could be a connection that actually dropped, and a single global threshold can't distinguish the two.

AWS Connect's agentic voice guidance recommends playing a holding prompt during lookup-style delays while keeping barge-in enabled throughout. The goal isn't eliminating silence — it's knowing in advance which stretches will run long and handling only those differently.

From Design to Operations: A Dead-Air Detection and Recovery Checklist

(a) Set the acceptance bar as numbers before writing code. Use FutureAGI's 3-second default as the starting alert line for a single gap, but target a tighter 15% total dead-air ratio per call rather than the 20% default. Declare a recovery-success rate (call continues without hanging up after a gap) of 85% or higher, and cap the delay before an unrecovered gap triggers human handoff at 5 seconds. These numbers are a starting point to be re-split by layer once your own call data accumulates.

(b) Failure patterns differ by layer: an RTP stream drops while the application layer assumes the call is still live; VAD endpointing waits several extra seconds after speech has clearly ended; STT emits only partial transcripts with no finalize event, so the next turn never starts; or a tool call or TTS generation/playback stalls with nothing going out over the line.

Recovery branches by the layer at fault. A telephony-layer drop ends the session immediately and logs a redial event; VAD/STT delay gets bridged with a short filler line; tool-call or TTS delay gets AWS's recommended holding prompt with barge-in left on. If a single call crosses the alert line twice, it escalates to human handoff automatically, with no approval step in the loop.

(c) The operations checklist starts with turn-level logging: store telephony, VAD, STT-finalize, LLM-first-token, tool-execution, TTS-first-audio, and playback timestamps under one turn ID so any gap can be traced back to its layer after the fact. Before shipping, run scenario tests that artificially inject tool-call and TTS-generation delay to confirm the holding prompt and handoff gate actually fire.

Repeated holding prompts and filler lines degrade the experience, so cap filler at three uses per call before forcing a handoff. Mask any audio segment in the logs that may contain PII.

(d) After launch, review dead-air events grouped by layer every week and feed recurring culprits — a specific STT vendor, a specific tool API — back into endpointing and timeout tuning. Track whether the total dead-air ratio and recovery-success rate improve release over release as the gate's own performance metric.

Takeaways to Use Right Away

Dead air looks like one symptom but hides separate causes across telephony, VAD, STT, tool calls, and TTS, so a single global threshold can't sort them out. Set an alert line at a 3-second single gap and 15% total dead air, attribute causes through turn-level logs that join per-layer timestamps, bridge delay-prone stretches with holding prompts and barge-in, and escalate repeated gaps to human handoff — that combination shrinks the window where silence turns into an abandoned call.

References

Voice Agent Dead Air Detection: Root Causes and Fixes — Hamming

Dead Air Detection — FutureAGI Docs

Agentic voice best practices — AWS Connect Admin Guide

Ask AI about this article

The assistant has read this article. Ask anything — it answers from the text and says so when something isn't in it.

Loading the chat…