What AMD Sorts: Human, Machine, Fax, Unknown
When an outbound callbot dials out, the loop first needs to know whether a human or an answering machine picked up before it can decide what to do next. Twilio's AMD offers two detection modes. MachineDetection=Enable returns a result as soon as a determination is reached, narrowing the outcome to four values — machine_start, human, fax, unknown — and fits predictive-dialer use cases built to cut agent idle time. DetectMessageEnd waits until the voicemail greeting actually ends (the beep), splitting the outcome into six values — machine_end_beep, machine_end_silence, machine_end_other, human, fax, unknown.
The real difference between the two modes isn't when detection stops, it's how long it waits. Synchronous AMD connects the callee only after a determination is made; asynchronous AMD (AsyncAmd=true) connects immediately and runs detection in the background, posting the result to a callback URL. Either way, letting an unknown result pass through without a retry leaves a blind spot: the script keeps running without ever knowing whether it reached a person or a mailbox.
Four Forked Streams, One Spent by Async AMD
While async AMD is running, it consumes one of the four forked streams allowed per call. That cap is shared with Media Streams, SIPREC, and Real-Time Transcription, so a call that already has real-time transcription running leaves only two streams of headroom once async AMD is added. Stack recording, analytics, and AMD onto the same call without tracking the cap, and whichever feature you add last gets silently rejected.
Speed and accuracy trade off against each other. A shorter detection timeout means less audio to analyze and more false positives; a longer one means slower responses and longer agent wait times. Switching to async mode hides the perceived latency, but it doesn't make the forked-stream cap disappear — if you're planning to scale concurrent calls, recalculate your stream budget before adopting async AMD, not after.
Wiring an AMD Gate Into the Dialer: From Design to Operations
Building an AMD gate starts with declaring three numbers up front: cap false machine-classification of a real human at 2% or under, cap playback delay after beep detection at 300ms, and cap forked-stream usage per call at 3 — 75% of the 4-stream limit. Without that third number, it takes far too long to trace why some other feature, like real-time transcription or Media Streams, is quietly getting rejected because of AMD.
Failures tend to repeat in the same shapes. First, shortening the timeout to speed up response and ending up misclassifying a real human as a machine, cutting the call early. Second, using Enable mode alone without DetectMessageEnd, so playback starts before the beep finishes and the first part of the message gets clipped.
Third, writing logic that assumes synchronous mode on a path — the Participants API, or <Dial><Number> — where async AMD is the fixed, non-configurable default. Fourth, never budgeting the forked-stream cap and stacking real-time transcription onto the same call, so either AMD or the transcript silently drops.
Declare the recovery branches ahead of time. Don't let an unknown result pass silently — route it to a human agent queue or re-probe with a short confirmation question — and when the forked-stream cap is reached, drop a lower-priority feature (recorded transcription, for instance) before you let AMD itself get squeezed out.
Before deploying, run a regression pass against a mixed sample of human voice, voicemail greetings, and fax tones, and check whether beep timing drifts because greeting length varies by carrier. Log CallSid, AnsweredBy, and MachineDetectionDuration as required fields, and keep consent and do-not-call records alongside them so pre-recorded playback never drifts outside policy.
If a pacing gate decides when to hang up, an AMD gate sits upstream deciding who to connect to in the first place. Review the false-positive rate and forked-stream exhaustion frequency weekly, retune the timeout and the share of calls running async AMD, and keep per-carrier baselines separate so a slow-beep pattern on one carrier doesn't get buried in the average.
Takeaways at a Glance
AMD isn't a simple human-or-machine filter — it's another in-call feature competing for the finite resource of forked streams. Hold the line at three numbers — false machine-classification under 2%, beep-to-playback delay under 300ms, stream usage under 3 — and pre-declare recovery branches for unknown results and resource contention alike, and the same gate holds up even when the carrier changes underneath it.
References
Answering Machine Detection — Twilio Docs
Answering Machine Detection FAQ & Best Practices — Twilio Docs
Ask AI about this article
The assistant has read this article. Ask anything — it answers from the text and says so when something isn't in it.
Loading the chat…