AI Agents for Business — Production Architecture for Custom AI Agent Development
Search interest has moved from “what is AI” to “AI agents for business”. Companies are no longer asking for a demo chatbot. They want custom AI agents that answer from their own data, take actions in their CRM and helpdesk, pick up the phone, and stay inside GDPR and the EU AI Act. This brief lays out the reference architecture stackcone (an AI agent development company) uses for custom AI agent development: one agent core with four pluggable layers — knowledge (RAG), action (tool calling + MCP), channels (web, Slack/Teams, WhatsApp, voice), and trust (evals, guardrails, human-in-the-loop).
Proposed outcome: A production AI agent that resolves 40–70% of routine requests end to end (support tickets, bookings, lead qualification, internal questions), cites its sources, writes back to your systems through audited tools, escalates to a human when unsure, and can be hosted in the EU or your own cloud.
What buyers are actually asking for
US Google data through mid-2026 shows agent and automation queries now pull roughly four times the volume of generic “AI for [task]” searches. “Agentic AI” sits around 110k monthly US searches, “AI agents” around 60k, and “AI agents for business” more than tripled year over year. The searches that turn into projects are more specific:
| Buyer | What they search | What they mean |
|---|---|---|
| SMB owner (US) | AI receptionist, AI voice agent, AI phone agent | Answer every call, book appointments, text a link, never miss a lead |
| SMB / ops lead (UK) | AI automation agency, n8n AI agent, AI workflow automation | Automate email, forms, invoices, CRM updates with LLM steps they can see |
| Support leader | AI customer support agent, AI agent for customer service | Deflect tickets in Zendesk/Intercom with grounded answers and real actions (refunds, order status) |
| Sales leader | AI sales agent, AI SDR, AI lead qualification | Qualify inbound leads, enrich, book meetings, update HubSpot/Salesforce |
| Knowledge-heavy firm | RAG chatbot for business, AI assistant for company documents | Ask questions across SharePoint, Drive, Confluence with citations and permissions |
| CTO / product team | custom AI agent development, MCP server development, LangGraph developer, multi-agent system | A partner who can ship a maintainable agent inside their product and cloud |
| EU / regulated buyer | GDPR compliant AI chatbot, EU AI Act compliant AI, on-premise LLM, private AI | Same as above, but data never leaves the EU and every decision is auditable |
Two patterns stand out. First, industry-specific agents (dental, legal, insurance, real estate, e-commerce) are the fastest-growing segment and outperform general-purpose ones. Second, roughly half of B2B buyers now start vendor research in ChatGPT, Claude or Perplexity rather than Google, and they ask about architecture, pricing and compliance directly — which is what this brief answers.
Why most AI agent projects stall
- Demo ≠ production — a prompt in a playground works on ten questions and fails on the eleventh; nobody measured accuracy before launch
- No grounding — the agent answers from model memory instead of the company’s current policies, prices, and documents, so it hallucinates
- Read-only agents — it can explain the refund policy but cannot issue the refund, so the ticket still reaches a human
- Unsafe actions — the opposite failure: an agent with broad API keys and no approval step, which security teams (rightly) block
- Framework lock-in — integrations written once for one vendor’s SDK have to be rewritten when the model or framework changes
- Compliance afterthought — EU customers ask where data is processed, how long logs are kept, and who reviews automated decisions; the answer is “we’ll check”
- No owner after launch — prompts drift, documents go stale, costs creep, and nobody watches the dashboard
Requirements
Functional
- Grounded answers: Retrieve from company sources (docs, website, tickets, product DB) with citations; refuse when the answer is not in the sources
- Actions through tools: Typed, permission-scoped tools for CRM, helpdesk, calendar, billing, e-commerce, and internal APIs
- Multi-channel: Website widget, Slack/Teams, WhatsApp/SMS, email, and phone (voice) from one agent core
- Memory: Conversation state per user/session; optional long-term customer profile from the CRM rather than free-form model memory
- Human-in-the-loop: Approval queue for risky actions (refund > threshold, contract changes); warm handoff to a human with full context
- Admin console: Edit knowledge sources, tool permissions, prompts and escalation rules without a code deploy
Non-functional
- Accuracy gate: Offline eval suite (golden questions + tool-call scenarios) must pass before every release
- Latency: First token < 1.5 s for chat; < 800 ms turn latency for voice
- Security: SSO, RBAC, per-tenant data isolation, secrets in a vault, least-privilege tool credentials
- Data residency: Selectable US / EU / UK regions; option for self-hosted open-weight models
- Observability: Full trace of every run — retrieved chunks, tool calls, tokens, cost, latency, user feedback
- Model independence: Swap between Claude, GPT, Gemini or open-weight models through one gateway without rewriting the agent
Reference architecture
The design keeps one agent core (planner + tool loop) and makes everything around it pluggable. Channels are thin adapters. Knowledge and tools sit behind stable interfaces — retrieval as a tool, integrations as MCP servers — so the same agent can run behind a web widget today and a phone line next quarter.
Reference architecture — thin channel adapters, one agent core, retrieval and MCP tools behind stable interfaces, and a trust layer (guardrails, approvals, tracing) around every run
Typical support-agent run — look up the record, ground the answer in policy, act within limits, and hand off with a draft when the action is risky
Build path
Ship one narrow agent to production first; add channels and tools only after the eval gate is green
Five agent patterns (and when to use each)
| Pattern | Typical use case | Core pieces | Complexity |
|---|---|---|---|
| 1. Knowledge agent (RAG) | Internal knowledge base, policy Q&A, product docs assistant | Ingestion, hybrid search, reranker, citations, ACL filtering | Low–medium |
| 2. Action agent (tool calling) | AI customer support agent, order status, refunds, CRM updates | RAG + typed tools via MCP, approval thresholds, audit log | Medium |
| 3. Voice agent | AI receptionist, appointment booking, lead qualification calls | Twilio/SIP, streaming STT/TTS (Vapi, Retell, LiveKit), short tool calls, SMS follow-up | Medium |
| 4. Workflow agent | Invoice/document processing, email triage, lead enrichment | n8n or Temporal workflow with LLM steps, structured outputs, retries | Low–medium |
| 5. Multi-agent system | Research + drafting pipelines, complex ops spanning many systems | Supervisor + specialist agents (LangGraph), shared state, per-agent permissions | High |
Rule of thumb: start at the lowest pattern that solves the problem. Most “we need a multi-agent system” requests are really pattern 2 with good retrieval. Move to pattern 5 only when one agent’s tool list, permissions, or prompt become unmanageable.
Indicative mix of inbound AI agent requests by pattern (stackcone pipeline, 2026)
Recommended stack
Recommendation: Python (FastAPI) agent service on LangGraph or a vendor agent SDK (Claude Agent SDK / OpenAI Agents SDK), Postgres + pgvector for state and retrieval, integrations exposed as MCP servers, a model gateway for provider switching, and Langfuse for tracing and evals. For SMB workflow agents, n8n (self-hosted in the EU if needed) keeps the automation visible to the client.
| Layer | Technology | Why |
|---|---|---|
| Agent runtime | LangGraph, Claude Agent SDK, or OpenAI Agents SDK | Explicit state graph, retries, human-interrupt nodes; SDKs are faster for single-agent builds |
| Models | Claude, GPT, Gemini; Llama/Mistral/Qwen for self-hosted | Route by task: small fast model for classification, frontier model for planning |
| Model gateway | LiteLLM or cloud-native (Bedrock, Azure OpenAI, Vertex) | One API, regional endpoints, fallback, per-tenant cost caps |
| Retrieval | pgvector or OpenSearch + BM25 hybrid, Cohere/Voyage reranker | Hybrid + rerank beats pure vector search on business documents |
| Integrations | MCP servers (TypeScript or Python) | Reusable across agents, IDEs and chat clients; clear per-tool scopes |
| Voice | Vapi or Retell (fast), LiveKit Agents (custom), Twilio numbers | Sub-second turn-taking and barge-in are hard to build from scratch |
| Workflows | n8n (SMB), Temporal (engineering teams) | Durable retries and schedules outside the LLM loop |
| Observability & evals | Langfuse or LangSmith, plus a pytest eval suite in CI | Trace every run; block releases that regress accuracy |
| Frontend | Next.js widget + admin console; Slack/Teams apps | Streaming responses (SSE), feedback buttons, source previews |
| Hosting | AWS / GCP / Azure in the client’s account and chosen region | Data residency and procurement are simpler when it runs in their cloud |
Why MCP for integrations? A HubSpot or Zendesk MCP server built once can be used by the support agent, the sales agent, and staff inside Claude or ChatGPT. It also forces clean tool contracts — name, typed arguments, scope — which is exactly what security reviews ask for.
Why not start with a no-code agent builder? They are fine for prototypes. Production buyers usually need custom retrieval, permission-aware data, eval gates and hosting control, which is where a custom build pays off. A common path is prototype in n8n, then move the core loop to code.
Typical monthly run-cost split for a mid-volume support agent (illustrative)
Component design
1 — Knowledge layer (RAG)
- Ingestion: Connectors for Drive, SharePoint, Confluence, Notion, website, and resolved tickets; scheduled re-sync with change detection
- Chunking: Structure-aware (headings, tables) with parent-document linking; metadata for source, owner, date, and access group
- Search: Hybrid BM25 + vector, reranker, top-k 5–8; filter by the caller’s access group before ranking
- Answering: Cite chunk IDs; say “I don’t know” when retrieval score is below threshold
- QA gate: Recall@5 and faithfulness on a 100–300 question golden set (see how to evaluate RAG retrieval)
2 — Action layer (tools + MCP)
- Tool contract: Name, description, JSON schema, scope (
read/write/money), and idempotency key - Permissions: Tools run with the end user’s or a service account’s least-privilege token, never an admin key
- Thresholds: Write actions above a configurable limit go to the approval queue
- Audit: Every call logged with input, output, user, and trace ID
- QA gate: Scenario tests that assert the right tool is called with the right arguments — and that forbidden tools are never called
3 — Agent core
- Loop: Plan → call tools → observe → answer, with a max-step budget and timeout
- State: Session state in Postgres; customer facts pulled from CRM on demand instead of stored in model memory
- Routing: Small model classifies intent and picks the playbook; frontier model handles multi-step reasoning
- Fallback: Provider failover through the gateway; graceful “I’ll pass this to a colleague” when all else fails
4 — Channels
| Channel | Adapter | Notes |
|---|---|---|
| Web chat | Next.js widget over SSE | Streaming, source cards, thumbs up/down |
| Slack / Teams | Bot app + events API | Uses the employee’s identity for permission-aware answers |
| WhatsApp / SMS | Twilio or WhatsApp Cloud API | Template messages for outbound; opt-in tracking |
| Voice | Vapi / Retell / LiveKit on Twilio numbers | Short tool calls, SMS follow-up for links, recording disclosure |
| Inbound parse webhook | Draft-for-approval by default |
5 — Trust layer
- Input: PII detection and redaction before logs; prompt-injection filtering on retrieved content and user input
- Output: Policy checks (no legal/medical advice beyond scope, no price promises outside catalogue), citation required for factual claims
- Human-in-the-loop: Approval inbox, handoff with summary and draft, confidence thresholds per intent
GDPR, UK GDPR & EU AI Act
European and UK buyers search for “GDPR compliant AI chatbot” and “EU AI Act compliant AI” because procurement will ask. The architecture above supports this by design:
| Requirement | How the design handles it |
|---|---|
| Data residency | In-region model endpoints (Bedrock, Azure OpenAI or Vertex AI in the required region) or self-hosted open-weight models; vector DB and logs in the same region |
| No training on customer data | Enterprise/API terms with zero data retention where offered; documented in the DPA |
| Data minimisation | PII redacted before tracing; configurable log retention (e.g. 30 days) |
| Right of access / erasure | Conversations keyed by user ID; delete endpoint cascades to logs and memory |
| Transparency (AI Act) | Clear “you are talking to an AI” disclosure in chat and at the start of voice calls |
| Human oversight | Approval queue and handoff for consequential decisions; no fully automated decisions with legal effect |
| Record keeping | Trace store with model version, prompt version, tools called, and reviewer actions |
| Calls & messaging | Recording disclosure, PECR/TCPA-aware outbound rules, opt-out handling for SMS/WhatsApp |
This is engineering guidance, not legal advice — the client’s DPO or counsel should confirm risk classification for their specific use case.
Implementation plan
Phase 1 — Discovery & golden set (week 1)
Pick one use case with clear volume and ROI (e.g. top 20 support intents). Collect 100–300 real questions and expected answers/actions. Map systems, APIs, permissions, and data residency needs. Agree success metrics: resolution rate, accuracy, CSAT, cost per conversation.
Risk: No API access to key systems — confirm credentials and sandbox accounts in week 1. Rollback: none needed; discovery output stands alone as a spec.
Phase 2 — Knowledge layer (week 2)
Ingest sources with access metadata, build hybrid search + reranker, and measure recall on the golden set. Fix chunking and missing content before touching prompts.
Risk: Stale or contradictory documents — surface conflicts to content owners. Rollback: read-only; no production impact.
Phase 3 — Tools & agent core (weeks 2–3)
Build 2–4 MCP tools (read first, then one write tool with thresholds). Wire the agent loop, state, and model gateway. Add scenario tests for every tool path.
Risk: Over-broad API scopes — create dedicated service accounts. Rollback: write tools behind a feature flag.
Phase 4 — Channel, guardrails & eval gate (weeks 3–4)
Ship the first channel (usually web widget or helpdesk sidebar), PII redaction, injection filtering, handoff flow, and CI eval suite. Release only when accuracy and tool-call tests pass agreed thresholds.
Risk: Accuracy plateau — expand golden set and fix retrieval before prompt tuning. Rollback: widget toggle; traffic back to humans.
Phase 5 — Shadow mode & go-live (weeks 4–5)
Agent drafts replies that humans approve for 1–2 weeks; compare against human answers. Then turn on auto-resolve for intents that clear the bar, keeping approval for the rest.
Risk: Edge cases not in golden set — feed every human correction back into evals. Rollback: per-intent switch back to draft-only.
Phase 6 — Expand channels & hand over (week 6+)
Add voice, Slack/Teams or WhatsApp on the same core; add more tools. Deliver runbook, dashboards, prompt/version registry, and admin training.
Risk: Cost growth with volume — per-tenant caps and model routing. Rollback: channels are independent adapters; disable one without affecting others.
Evals, monitoring & ops
| Signal | Source | Cadence |
|---|---|---|
| Resolution rate (no human needed) | Conversation outcomes | Daily dashboard; weekly review |
| Answer faithfulness & citation rate | LLM-judge + sampled human review | Every release (CI) + weekly sample |
| Tool-call correctness | Scenario test suite | Every release |
| Handoff rate by intent | Trace store | Weekly — target intents above 30% for improvement |
| Latency p50 / p95 | Tracing | Real-time alerts |
| Cost per conversation | Gateway token + tool usage | Weekly; alert on 30% jumps |
| CSAT / thumbs feedback | Widget + helpdesk | Weekly |
| Knowledge freshness | Ingestion job status | Daily; alert on failed syncs |
More on this in AI agent testing and evaluation strategies.
Indicative auto-resolution ramp — shadow mode first, then per-intent enablement (illustrative)
Cost & effort estimate
| Scope | Timeline | Indicative cost (USD; GBP/EUR similar) |
|---|---|---|
| Pilot: 1 use case, 1 channel, RAG + 2–4 tools, eval gate | 3–6 weeks | 8k–25k |
| AI receptionist / voice agent with booking + CRM | 3–5 weeks | 10k–30k |
| Production multi-channel agent with SSO, admin console, EU hosting | 8–14 weeks | 25k–80k |
| Multi-agent system across several departments | 12+ weeks | 60k+ |
| Run costs (models, hosting, voice minutes, tracing) | monthly | 300–3,000 depending on volume |
| Tuning retainer (evals, prompts, new intents) | monthly | 1k–4k |
Ranges depend on API quality of existing systems, number of integrations, compliance scope, and how clean the source documents are.
FAQ
How much does it cost to build a custom AI agent for a business?
A focused pilot usually costs 8k–25k over 3–6 weeks. Multi-channel production agents with voice, evals and SSO land in the 25k–80k range, plus monthly run and tuning costs.
What is the difference between an AI chatbot and an AI agent?
A chatbot answers. An agent also acts — it calls your CRM, helpdesk, calendar or billing APIs, follows multi-step plans, and hands off to a human when unsure.
Do I need a multi-agent system?
Usually not at first. One well-scoped agent with good retrieval and a few tools covers most use cases. Split into multiple agents when permissions or tool count make one agent unmanageable.
Can an AI agent be GDPR and EU AI Act compliant?
Yes — EU-region or self-hosted models, zero data retention, PII redaction, a DPA, AI disclosure, and human oversight for consequential decisions.
Who builds custom AI agents like this?
stackcone is an AI agent development company founded by Amar Kumar. We build RAG chatbots, AI customer support agents, AI voice agents and MCP integrations for businesses, with fixed-scope pilots and full handover. See our portfolio or contact us.
What is MCP and why does it matter?
The Model Context Protocol is an open standard for exposing tools and data to AI models. MCP-based integrations work across Claude, ChatGPT, IDEs and custom agents, so you are not locked into one framework. See AI agents and MCP in practice.
Glossary
| Term | Meaning |
|---|---|
| AI agent | An LLM-driven program that plans, calls tools, and acts toward a goal, not just answers text |
| Agentic AI | Umbrella term for systems where models take multi-step actions with some autonomy |
| RAG | Retrieval-augmented generation — answering from retrieved company documents with citations |
| Tool calling | The model returns a structured request to run a function (e.g. create_refund); your code executes it |
| MCP | Model Context Protocol — open standard for connecting models to tools and data sources |
| Human-in-the-loop | A person approves or edits agent actions above a risk threshold |
| Golden set | Curated real questions and expected answers/actions used to measure accuracy |
| Shadow mode | Agent drafts responses that humans review before anything reaches the customer |
| Model gateway | A single API in front of several LLM providers for routing, fallback, and cost control |