How to Build an AI Voice Agent: Technical Architecture
This brief is the technical blueprint for building a production AI voice agent — the kind that answers a real phone line, holds a natural conversation, calls your tools mid-call, and hands off to a person cleanly. It covers the full pipeline: telephony, streaming speech-to-text, the LLM agent loop, tool calling, streaming text-to-speech, the latency budget that makes or breaks the experience, and how to deploy and monitor it.
Who this is for: engineers and technical founders scoping a voice agent build — whether on a managed platform (Vapi, Retell, LiveKit) or a custom pipeline — who want the real architecture, not a marketing page.
How a voice agent actually works
A phone call is an audio stream in both directions. A voice agent turns the caller's speech into text, decides what to say or do, and turns the reply back into speech — fast enough that it feels like a conversation, not a transaction with a machine.
Four capabilities have to work together in real time:
- Hear: transcribe speech to text as the caller talks (streaming, not wait-for-silence)
- Think: an LLM agent decides what to say and whether to call a tool
- Act: call your systems — calendar, CRM, order lookup — mid-conversation
- Speak: turn the reply into natural audio and stream it back, interruptible
None of this is exotic technology on its own. The engineering problem is latency and turn-taking: every stage adds delay, callers notice a pause over roughly a second, and the agent has to know when the caller is done talking versus just pausing to think.
Reference architecture
Reference architecture — every stage streams; the orchestrator owns turn-taking, barge-in, and when to call a tool vs keep talking
TTS stops immediately (barge-in)
One conversation turn — transcription, tool call, and streamed reply, target under ~800ms end to end
Call lifecycle
The loop (listen → decide → act → speak) repeats each turn until the call resolves or hands off to a human
Indicative latency budget per stage — target ~700-900ms total turn time
The five pipeline stages
1 — Telephony
- Number: Twilio, Telnyx, or SIP trunk into your existing PBX
- Media streaming: Twilio Media Streams (or equivalent) opens a WebSocket carrying raw audio chunks both ways — this is what lets you process audio in real time instead of waiting for a recorded file
- Inbound and outbound: inbound answers calls to a number; outbound dials from your system (reminders, follow-ups) — outbound calling to consumers has stricter consent rules (TCPA in the US, PECR in the UK/EU)
2 — Streaming speech-to-text (STT)
- Providers: Deepgram, AssemblyAI, Azure Speech, or Whisper-based streaming services
- Partial vs final transcripts: partials update as the caller talks (for the agent to start "thinking early"); a final transcript fires on endpoint detection (the system's best guess that the caller finished a thought)
- Endpointing tuning: too aggressive and the agent interrupts; too relaxed and replies feel slow. This is one of the highest-leverage tuning knobs in the whole system
3 — Agent loop (the LLM)
- Model choice: a fast, low-latency model (e.g. a smaller Claude, GPT, or Gemini variant) for turn-taking; escalate to a stronger model only for complex reasoning steps
- Prompt: persona, business info, conversation rules, and available tools — kept short, since every token adds latency
- State: conversation history, caller identity (from caller ID or earlier verification), and call metadata
- Streaming output: the LLM's tokens stream directly into TTS as they're generated, rather than waiting for the full reply
4 — Tool calling
- Typed functions:
check_availability,book_appointment,lookup_order,send_sms,transfer_call— same pattern as any AI agent's tools - Filler phrases: if a tool call takes more than ~500ms, the agent says something ("let me check that for you") so the line doesn't go silent
- Timeouts: every tool call needs a hard timeout and a fallback ("I'm having trouble pulling that up — let me have someone call you back")
5 — Streaming text-to-speech (TTS)
- Providers: ElevenLabs, Cartesia, PlayHT, or OpenAI's voice models
- Streaming: audio starts playing from the first generated chunk, not after the full sentence is synthesized
- Voice selection: match tone to the business; disclose that it's an AI voice per your jurisdiction's rules
Latency budget
People notice a pause above roughly 1 second and start talking over the agent or assuming the call dropped. A well-tuned agent keeps the full turn — from the caller finishing a sentence to the agent's audio starting — under about 700–900ms. That budget is split roughly like this:
| Stage | Typical latency | How to reduce it |
|---|---|---|
| Endpoint detection (STT) | 100–300ms | Tune endpointing sensitivity; use a provider with fast finalization |
| LLM time-to-first-token | 150–400ms | Smaller/faster model, short prompt, prompt caching |
| Tool call (if any) | 100–500ms | Cache lookups, async where possible, filler phrase if slow |
| TTS time-to-first-audio | 150–300ms | Streaming synthesis provider, shorter first sentence |
| Network (telephony + media) | 50–150ms | Co-locate services with your telephony provider's region |
The biggest win is almost always streaming every stage rather than waiting for each one to fully finish before starting the next — transcribe while the caller talks, generate LLM tokens while streaming them to TTS, and start playing audio from the first synthesized chunk.
Tool calling mid-call
This is what separates a voice agent from an IVR menu: it can actually do something. The pattern is the same structured tool-calling used by any AI agent — the LLM returns a typed function call instead of text, your backend executes it against a real system, and the result feeds back into the conversation.
| Tool | Typical backend | Notes |
|---|---|---|
check_availability / book_appointment | Calendar API, booking software (Calendly, ServiceM8, practice management) | Return the nearest 2-3 open slots, not a full calendar dump |
lookup_order / lookup_account | CRM, order database, helpdesk | Verify identity first (name + one more field) |
send_sms | Twilio Messaging | For anything easier to read than to hear — links, addresses, confirmation codes |
transfer_call | Telephony provider | Warm transfer with a spoken summary to the human, not a cold handoff |
create_ticket | Helpdesk / CRM | For anything the agent can't resolve on the call |
Interruptions and barge-in
Real conversations have interruptions. A voice agent that can't be talked over feels robotic and frustrates callers who already know what they want.
- Continuous listening: voice activity detection runs on the input stream even while the agent's own audio is playing
- Immediate stop: the moment the caller starts speaking, TTS playback is cut and the unspoken remainder is discarded
- Context repair: the agent should acknowledge the interruption naturally ("sorry, go ahead") rather than restarting its previous sentence
- False positives: background noise or a cough can trigger false barge-in; noise suppression and a short confirmation window reduce this
Managed voice platforms (Vapi, Retell, LiveKit) handle this out of the box. A custom pipeline needs an explicit interrupt signal wired from the input VAD into the TTS output stream — this is one of the more fiddly parts of building from scratch.
Managed platform vs custom build
| Option | Best for | Trade-off |
|---|---|---|
| Vapi | Fast launch, flexible tool calling, bring-your-own models | You design the conversation flows; multi-provider costs add up |
| Retell AI | Fast launch, strong call quality out of the box | Less low-level control |
| LiveKit Agents (open source) | Full control, self-hosting, data residency requirements | You own the telephony/media reliability and scaling |
| Fully custom (Twilio Media Streams + your own STT/LLM/TTS wiring) | Very high volume, unusual latency or compliance requirements | Months of engineering on barge-in, endpointing, and failover before it's production-grade |
For most businesses, Vapi or Retell gets a production-grade voice agent live in 3-5 weeks. Build custom only when you have a specific reason — self-hosting for data residency, extreme call volume, or telephony features those platforms don't expose. The work that actually differentiates a good voice agent — the script, the tools, the escalation rules, the testing — is the same either way.
Building one? stackcone builds production voice agents on Vapi, Retell and LiveKit — telephony setup, tool integrations, latency tuning and call testing included. Talk to stackcone →
Recommended stack
| Layer | Technology | Why |
|---|---|---|
| Telephony | Twilio (or Telnyx) | Reliable Media Streams API, global numbers, good docs |
| Voice platform | Vapi or Retell AI | Bundles STT, LLM orchestration, TTS, barge-in, and tool calling behind one API |
| STT (if unbundled) | Deepgram or AssemblyAI streaming | Low-latency partial transcripts, strong accuracy on phone audio |
| LLM | Fast model tier (Claude, GPT, or Gemini "flash"/"mini" class) | Time-to-first-token matters more than raw reasoning depth for most calls |
| TTS (if unbundled) | ElevenLabs or Cartesia streaming | Natural voice with low time-to-first-audio |
| Tools | MCP servers or typed REST functions | Reusable across this and other agents (see our AI agents architecture brief) |
| Backend | Python (FastAPI) or Node | Webhook handlers for tool calls, booking logic, CRM sync |
| Observability | Call recordings + transcripts + Langfuse/LangSmith traces | Every call reviewable; latency and tool-call failures visible |
Implementation plan
Phase 1 — Script & discovery (week 1)
Listen to 30-50 real calls (or write realistic scenarios for a new line). List the top reasons people call and the right outcome for each. Choose the platform (Vapi/Retell vs custom) and provision a Twilio number.
Risk: no call recordings to learn from — run a short pilot with a human answering and recording first. Rollback: none needed; discovery is standalone.
Phase 2 — Core conversation (week 2)
Build the greeting, FAQ handling, and identity verification. Tune STT endpointing and TTS voice. Get the basic back-and-forth feeling natural before adding tools.
Risk: agent talks over callers or feels sluggish — this phase is where most latency and barge-in tuning happens. Rollback: keep the number pointed at a human fallback until this phase passes internal testing.
Phase 3 — Tool calling (weeks 2-3)
Wire 2-4 tools: booking/calendar, lookup, SMS, transfer. Add filler phrases and timeouts. Test each tool path with deliberately slow and failing responses.
Risk: a hanging tool call goes silent on the line — enforce timeouts everywhere. Rollback: tools behind a feature flag; agent falls back to "I'll have someone call you back."
Phase 4 — Testing & hardening (week 4)
Run 50+ test calls: accents, background noise, interruptions, edge cases, angry callers, silence. Fix the script and endpointing based on real transcripts. Add call recording and disclosure.
Risk: works in quiet testing, breaks on noisy mobile calls — test on real phones, not just a browser mic. Rollback: soft-launch on a secondary number before porting the main line.
Phase 5 — Go-live & monitoring (week 5)
Port or forward the production number. Daily transcript review for the first two weeks. Track resolution rate, transfer rate, and latency percentiles.
Risk: volume spikes reveal issues testing didn't catch — start with after-hours-only traffic if you can. Rollback: instant forward back to human answering.
Testing and monitoring
| Signal | Source | Cadence |
|---|---|---|
| Turn latency (p50/p95) | Voice platform dashboard or custom tracing | Real-time alerting |
| Call resolution rate | Call outcome tags | Weekly |
| Transfer / escalation rate | Call logs | Weekly — tune script if too high |
| Barge-in / interruption handling | Sampled transcript review | Weekly for first month |
| Tool-call success rate | Tool execution logs | Daily alert on failures |
| Caller sentiment (sampled) | Transcript review or post-call SMS survey | Weekly |
For the accuracy and regression-testing side of the agent itself (not just the audio pipeline), see AI agent testing and evaluation strategies.
Cost & effort estimate
| Scope | Timeline | Indicative cost (USD) |
|---|---|---|
| Voice agent on a managed platform (Vapi/Retell), 2-4 tools | 3-5 weeks | $10k–30k |
| Custom pipeline (self-hosted STT/LLM/TTS, LiveKit) | 8-12 weeks | $30k–70k |
| Run cost per minute (platform + telephony + speech + LLM) | monthly | $0.08–0.30/min |
| Ongoing tuning (script, endpointing, new tools) | monthly | $500–2,000 |
FAQ
What is the fastest way to build an AI voice agent?
Use a managed voice platform (Vapi, Retell, or Bland) that bundles telephony, streaming STT/TTS, and tool calling behind one API. Configure a prompt, connect tools, and you have a working agent in days. Build the pipeline yourself only when you need self-hosting, unusual latency budgets, or telephony features those platforms don't expose.
What causes latency in a voice agent and how do you fix it?
The main sources are STT finalization, LLM time-to-first-token, tool calls, and TTS generation. Fix it by streaming every stage, using a fast model for turn-taking, pre-fetching data where possible, and keeping tool calls under 300ms with cached or async fallbacks.
How do you handle interruptions (barge-in)?
Run voice activity detection continuously, even while the agent speaks. When the caller starts talking, stop TTS playback immediately and feed the new input to the agent. Most voice platforms handle this natively; a custom build needs an explicit interrupt signal.
Can a voice agent call functions or APIs mid-call?
Yes — the same structured tool-calling used by any AI agent. The LLM returns a typed function call, your backend executes it, and the result feeds back into the conversation. Keep calls fast and give the agent a filler phrase for anything slower than ~500ms.
Who builds custom AI voice agents?
stackcone builds production AI voice agents on Vapi, Retell and LiveKit for businesses, including telephony setup, tool integrations, latency tuning, and call testing. See our AI voice agents for business guide for the buyer-side view, or contact us to scope a build.
Glossary
| Term | Meaning |
|---|---|
| STT / ASR | Speech-to-text / automatic speech recognition — converts audio to text |
| TTS | Text-to-speech — converts text to natural audio |
| Endpointing | Detecting that the caller has finished a thought and it's the agent's turn |
| Barge-in | The caller interrupting the agent while it's speaking; the agent must stop and listen |
| Time-to-first-token / -audio | How long until the LLM's or TTS's first output chunk is ready — the key latency metric, not total generation time |
| Media Streams | Twilio's WebSocket API that streams raw call audio in real time, instead of a recorded file |
| Warm transfer | Transferring a call to a human with a spoken summary, rather than a cold, contextless handoff |
| Voice activity detection (VAD) | Detecting when someone is speaking versus silence or background noise |