How to Build an AI Voice Agent: Technical Architecture

By Amar Kumar, founder of stackcone · September 2026

This brief is the technical blueprint for building a production AI voice agent — the kind that answers a real phone line, holds a natural conversation, calls your tools mid-call, and hands off to a person cleanly. It covers the full pipeline: telephony, streaming speech-to-text, the LLM agent loop, tool calling, streaming text-to-speech, the latency budget that makes or breaks the experience, and how to deploy and monitor it.

Who this is for: engineers and technical founders scoping a voice agent build — whether on a managed platform (Vapi, Retell, LiveKit) or a custom pipeline — who want the real architecture, not a marketing page.

How a voice agent actually works

A phone call is an audio stream in both directions. A voice agent turns the caller's speech into text, decides what to say or do, and turns the reply back into speech — fast enough that it feels like a conversation, not a transaction with a machine.

Four capabilities have to work together in real time:

None of this is exotic technology on its own. The engineering problem is latency and turn-taking: every stage adds delay, callers notice a pause over roughly a second, and the agent has to know when the caller is done talking versus just pausing to think.

Reference architecture

flowchart TB classDef tel fill:#ede9fe,stroke:#7c3aed,color:#5b21b6 classDef speech fill:#dbeafe,stroke:#2563eb,color:#1e3a8a classDef agent fill:#ccfbf1,stroke:#0d9488,color:#115e59 classDef tool fill:#f1f5f9,stroke:#64748b,color:#334155 classDef mon fill:#fef3c7,stroke:#d97706,color:#92400e CALLER["Caller\\nPSTN / mobile"]:::tel SIP["Telephony\\nTwilio / Telnyx / SIP trunk"]:::tel CALLER <-->|"audio stream"| SIP WS["Media WebSocket\\nbidirectional audio"]:::speech SIP <--> WS VAD["Voice activity detection\\n+ endpointing"]:::speech STT["Streaming STT\\nDeepgram / Whisper / Azure"]:::speech WS --> VAD --> STT ORCH["Voice orchestrator\\nturn-taking + state"]:::agent STT --> ORCH LLM["LLM agent loop\\nfast model, short context"]:::agent ORCH <--> LLM TOOLS["Tool calls\\ncalendar, CRM, order lookup, SMS"]:::tool LLM -->|"function call"| TOOLS TOOLS -->|"result"| LLM TTS["Streaming TTS\\nElevenLabs / Cartesia / OpenAI"]:::speech ORCH --> TTS TTS --> WS OBS["Call recording, transcript,\\nlatency + tool traces"]:::mon ORCH -.-> OBS LLM -.-> OBS

Reference architecture — every stage streams; the orchestrator owns turn-taking, barge-in, and when to call a tool vs keep talking

sequenceDiagram autonumber participant C as Caller participant T as Telephony (Twilio) participant S as Streaming STT participant A as Agent (LLM) participant F as Tool / API participant V as Streaming TTS C->>T: speaks ("book me Tuesday at 3") T->>S: audio chunks (continuous) S-->>A: partial transcript (interim) S->>A: final transcript (endpoint detected) A->>A: decide: reply or call tool A->>F: check_availability(tuesday, 3pm) F->>A: slot open A->>V: response tokens (streamed) V-->>T: audio chunks (streamed) T-->>C: hears reply, ~700ms after finishing speaking Note over A,F: If caller speaks again mid-reply,
TTS stops immediately (barge-in)

One conversation turn — transcription, tool call, and streamed reply, target under ~800ms end to end

Call lifecycle

The loop (listen → decide → act → speak) repeats each turn until the call resolves or hands off to a human

Indicative latency budget per stage — target ~700-900ms total turn time

The five pipeline stages

1 — Telephony

2 — Streaming speech-to-text (STT)

3 — Agent loop (the LLM)

4 — Tool calling

5 — Streaming text-to-speech (TTS)

Latency budget

People notice a pause above roughly 1 second and start talking over the agent or assuming the call dropped. A well-tuned agent keeps the full turn — from the caller finishing a sentence to the agent's audio starting — under about 700–900ms. That budget is split roughly like this:

StageTypical latencyHow to reduce it
Endpoint detection (STT)100–300msTune endpointing sensitivity; use a provider with fast finalization
LLM time-to-first-token150–400msSmaller/faster model, short prompt, prompt caching
Tool call (if any)100–500msCache lookups, async where possible, filler phrase if slow
TTS time-to-first-audio150–300msStreaming synthesis provider, shorter first sentence
Network (telephony + media)50–150msCo-locate services with your telephony provider's region

The biggest win is almost always streaming every stage rather than waiting for each one to fully finish before starting the next — transcribe while the caller talks, generate LLM tokens while streaming them to TTS, and start playing audio from the first synthesized chunk.

Tool calling mid-call

This is what separates a voice agent from an IVR menu: it can actually do something. The pattern is the same structured tool-calling used by any AI agent — the LLM returns a typed function call instead of text, your backend executes it against a real system, and the result feeds back into the conversation.

ToolTypical backendNotes
check_availability / book_appointmentCalendar API, booking software (Calendly, ServiceM8, practice management)Return the nearest 2-3 open slots, not a full calendar dump
lookup_order / lookup_accountCRM, order database, helpdeskVerify identity first (name + one more field)
send_smsTwilio MessagingFor anything easier to read than to hear — links, addresses, confirmation codes
transfer_callTelephony providerWarm transfer with a spoken summary to the human, not a cold handoff
create_ticketHelpdesk / CRMFor anything the agent can't resolve on the call

Interruptions and barge-in

Real conversations have interruptions. A voice agent that can't be talked over feels robotic and frustrates callers who already know what they want.

Managed voice platforms (Vapi, Retell, LiveKit) handle this out of the box. A custom pipeline needs an explicit interrupt signal wired from the input VAD into the TTS output stream — this is one of the more fiddly parts of building from scratch.

Managed platform vs custom build

OptionBest forTrade-off
VapiFast launch, flexible tool calling, bring-your-own modelsYou design the conversation flows; multi-provider costs add up
Retell AIFast launch, strong call quality out of the boxLess low-level control
LiveKit Agents (open source)Full control, self-hosting, data residency requirementsYou own the telephony/media reliability and scaling
Fully custom (Twilio Media Streams + your own STT/LLM/TTS wiring)Very high volume, unusual latency or compliance requirementsMonths of engineering on barge-in, endpointing, and failover before it's production-grade

For most businesses, Vapi or Retell gets a production-grade voice agent live in 3-5 weeks. Build custom only when you have a specific reason — self-hosting for data residency, extreme call volume, or telephony features those platforms don't expose. The work that actually differentiates a good voice agent — the script, the tools, the escalation rules, the testing — is the same either way.

Building one? stackcone builds production voice agents on Vapi, Retell and LiveKit — telephony setup, tool integrations, latency tuning and call testing included. Talk to stackcone →

LayerTechnologyWhy
TelephonyTwilio (or Telnyx)Reliable Media Streams API, global numbers, good docs
Voice platformVapi or Retell AIBundles STT, LLM orchestration, TTS, barge-in, and tool calling behind one API
STT (if unbundled)Deepgram or AssemblyAI streamingLow-latency partial transcripts, strong accuracy on phone audio
LLMFast model tier (Claude, GPT, or Gemini "flash"/"mini" class)Time-to-first-token matters more than raw reasoning depth for most calls
TTS (if unbundled)ElevenLabs or Cartesia streamingNatural voice with low time-to-first-audio
ToolsMCP servers or typed REST functionsReusable across this and other agents (see our AI agents architecture brief)
BackendPython (FastAPI) or NodeWebhook handlers for tool calls, booking logic, CRM sync
ObservabilityCall recordings + transcripts + Langfuse/LangSmith tracesEvery call reviewable; latency and tool-call failures visible

Implementation plan

Phase 1 — Script & discovery (week 1)

Listen to 30-50 real calls (or write realistic scenarios for a new line). List the top reasons people call and the right outcome for each. Choose the platform (Vapi/Retell vs custom) and provision a Twilio number.

Risk: no call recordings to learn from — run a short pilot with a human answering and recording first. Rollback: none needed; discovery is standalone.

Phase 2 — Core conversation (week 2)

Build the greeting, FAQ handling, and identity verification. Tune STT endpointing and TTS voice. Get the basic back-and-forth feeling natural before adding tools.

Risk: agent talks over callers or feels sluggish — this phase is where most latency and barge-in tuning happens. Rollback: keep the number pointed at a human fallback until this phase passes internal testing.

Phase 3 — Tool calling (weeks 2-3)

Wire 2-4 tools: booking/calendar, lookup, SMS, transfer. Add filler phrases and timeouts. Test each tool path with deliberately slow and failing responses.

Risk: a hanging tool call goes silent on the line — enforce timeouts everywhere. Rollback: tools behind a feature flag; agent falls back to "I'll have someone call you back."

Phase 4 — Testing & hardening (week 4)

Run 50+ test calls: accents, background noise, interruptions, edge cases, angry callers, silence. Fix the script and endpointing based on real transcripts. Add call recording and disclosure.

Risk: works in quiet testing, breaks on noisy mobile calls — test on real phones, not just a browser mic. Rollback: soft-launch on a secondary number before porting the main line.

Phase 5 — Go-live & monitoring (week 5)

Port or forward the production number. Daily transcript review for the first two weeks. Track resolution rate, transfer rate, and latency percentiles.

Risk: volume spikes reveal issues testing didn't catch — start with after-hours-only traffic if you can. Rollback: instant forward back to human answering.

Testing and monitoring

SignalSourceCadence
Turn latency (p50/p95)Voice platform dashboard or custom tracingReal-time alerting
Call resolution rateCall outcome tagsWeekly
Transfer / escalation rateCall logsWeekly — tune script if too high
Barge-in / interruption handlingSampled transcript reviewWeekly for first month
Tool-call success rateTool execution logsDaily alert on failures
Caller sentiment (sampled)Transcript review or post-call SMS surveyWeekly

For the accuracy and regression-testing side of the agent itself (not just the audio pipeline), see AI agent testing and evaluation strategies.

Cost & effort estimate

ScopeTimelineIndicative cost (USD)
Voice agent on a managed platform (Vapi/Retell), 2-4 tools3-5 weeks$10k–30k
Custom pipeline (self-hosted STT/LLM/TTS, LiveKit)8-12 weeks$30k–70k
Run cost per minute (platform + telephony + speech + LLM)monthly$0.08–0.30/min
Ongoing tuning (script, endpointing, new tools)monthly$500–2,000

FAQ

What is the fastest way to build an AI voice agent?

Use a managed voice platform (Vapi, Retell, or Bland) that bundles telephony, streaming STT/TTS, and tool calling behind one API. Configure a prompt, connect tools, and you have a working agent in days. Build the pipeline yourself only when you need self-hosting, unusual latency budgets, or telephony features those platforms don't expose.

What causes latency in a voice agent and how do you fix it?

The main sources are STT finalization, LLM time-to-first-token, tool calls, and TTS generation. Fix it by streaming every stage, using a fast model for turn-taking, pre-fetching data where possible, and keeping tool calls under 300ms with cached or async fallbacks.

How do you handle interruptions (barge-in)?

Run voice activity detection continuously, even while the agent speaks. When the caller starts talking, stop TTS playback immediately and feed the new input to the agent. Most voice platforms handle this natively; a custom build needs an explicit interrupt signal.

Can a voice agent call functions or APIs mid-call?

Yes — the same structured tool-calling used by any AI agent. The LLM returns a typed function call, your backend executes it, and the result feeds back into the conversation. Keep calls fast and give the agent a filler phrase for anything slower than ~500ms.

Who builds custom AI voice agents?

stackcone builds production AI voice agents on Vapi, Retell and LiveKit for businesses, including telephony setup, tool integrations, latency tuning, and call testing. See our AI voice agents for business guide for the buyer-side view, or contact us to scope a build.

Glossary

TermMeaning
STT / ASRSpeech-to-text / automatic speech recognition — converts audio to text
TTSText-to-speech — converts text to natural audio
EndpointingDetecting that the caller has finished a thought and it's the agent's turn
Barge-inThe caller interrupting the agent while it's speaking; the agent must stop and listen
Time-to-first-token / -audioHow long until the LLM's or TTS's first output chunk is ready — the key latency metric, not total generation time
Media StreamsTwilio's WebSocket API that streams raw call audio in real time, instead of a recorded file
Warm transferTransferring a call to a human with a spoken summary, rather than a cold, contextless handoff
Voice activity detection (VAD)Detecting when someone is speaking versus silence or background noise