AI Agents for Business — Production Architecture for Custom AI Agent Development

By Amar Kumar, founder of stackcone · September 2026

Search interest has moved from “what is AI” to “AI agents for business”. Companies are no longer asking for a demo chatbot. They want custom AI agents that answer from their own data, take actions in their CRM and helpdesk, pick up the phone, and stay inside GDPR and the EU AI Act. This brief lays out the reference architecture stackcone (an AI agent development company) uses for custom AI agent development: one agent core with four pluggable layers — knowledge (RAG), action (tool calling + MCP), channels (web, Slack/Teams, WhatsApp, voice), and trust (evals, guardrails, human-in-the-loop).

Proposed outcome: A production AI agent that resolves 40–70% of routine requests end to end (support tickets, bookings, lead qualification, internal questions), cites its sources, writes back to your systems through audited tools, escalates to a human when unsure, and can be hosted in the EU or your own cloud.

What buyers are actually asking for

US Google data through mid-2026 shows agent and automation queries now pull roughly four times the volume of generic “AI for [task]” searches. “Agentic AI” sits around 110k monthly US searches, “AI agents” around 60k, and “AI agents for business” more than tripled year over year. The searches that turn into projects are more specific:

BuyerWhat they searchWhat they mean
SMB owner (US)AI receptionist, AI voice agent, AI phone agentAnswer every call, book appointments, text a link, never miss a lead
SMB / ops lead (UK)AI automation agency, n8n AI agent, AI workflow automationAutomate email, forms, invoices, CRM updates with LLM steps they can see
Support leaderAI customer support agent, AI agent for customer serviceDeflect tickets in Zendesk/Intercom with grounded answers and real actions (refunds, order status)
Sales leaderAI sales agent, AI SDR, AI lead qualificationQualify inbound leads, enrich, book meetings, update HubSpot/Salesforce
Knowledge-heavy firmRAG chatbot for business, AI assistant for company documentsAsk questions across SharePoint, Drive, Confluence with citations and permissions
CTO / product teamcustom AI agent development, MCP server development, LangGraph developer, multi-agent systemA partner who can ship a maintainable agent inside their product and cloud
EU / regulated buyerGDPR compliant AI chatbot, EU AI Act compliant AI, on-premise LLM, private AISame as above, but data never leaves the EU and every decision is auditable

Two patterns stand out. First, industry-specific agents (dental, legal, insurance, real estate, e-commerce) are the fastest-growing segment and outperform general-purpose ones. Second, roughly half of B2B buyers now start vendor research in ChatGPT, Claude or Perplexity rather than Google, and they ask about architecture, pricing and compliance directly — which is what this brief answers.

Why most AI agent projects stall

Requirements

Functional

Non-functional

Reference architecture

The design keeps one agent core (planner + tool loop) and makes everything around it pluggable. Channels are thin adapters. Knowledge and tools sit behind stable interfaces — retrieval as a tool, integrations as MCP servers — so the same agent can run behind a web widget today and a phone line next quarter.

flowchart TB classDef chan fill:#ede9fe,stroke:#7c3aed,color:#5b21b6 classDef core fill:#dbeafe,stroke:#2563eb,color:#1e3a8a classDef know fill:#ccfbf1,stroke:#0d9488,color:#115e59 classDef act fill:#f1f5f9,stroke:#64748b,color:#334155 classDef trust fill:#fef3c7,stroke:#d97706,color:#92400e WEB["Website widget"]:::chan CHAT["Slack / Teams / WhatsApp"]:::chan VOICE["Phone: Twilio + voice runtime\nSTT / TTS"]:::chan API["Your product API"]:::chan GW["Channel gateway\nauth, session, rate limits"]:::core WEB --> GW CHAT --> GW VOICE --> GW API --> GW GR_IN["Input guardrails\nPII redaction, injection filter"]:::trust GW --> GR_IN AGENT["Agent core\nplanner + tool loop + state"]:::core GR_IN --> AGENT LLM["LLM gateway\nClaude / GPT / Gemini / open-weight"]:::core AGENT <--> LLM RAG["Retrieval tool\nhybrid search + rerank"]:::know VDB["Vector + keyword index\npgvector / OpenSearch"]:::know SRC["Sources: Drive, SharePoint,\nConfluence, site, tickets"]:::know AGENT --> RAG RAG --> VDB SRC -->|"ingest + ACLs"| VDB MCP["MCP servers\ntyped, scoped tools"]:::act AGENT --> MCP MCP --> CRM["CRM: HubSpot / Salesforce"]:::act MCP --> DESK["Helpdesk: Zendesk / Intercom"]:::act MCP --> OPS["Calendar, Stripe, Shopify,\ninternal APIs"]:::act HITL["Approval queue +\nhuman handoff"]:::trust AGENT -->|"risky action / low confidence"| HITL OBS["Tracing + evals\nLangfuse / LangSmith"]:::trust AGENT -.-> OBS

Reference architecture — thin channel adapters, one agent core, retrieval and MCP tools behind stable interfaces, and a trust layer (guardrails, approvals, tracing) around every run

sequenceDiagram autonumber participant U as Customer participant G as Gateway + guardrails participant A as Agent core participant R as Retrieval participant T as MCP tools participant H as Human agent U->>G: "Where is my order? I want a refund" G->>A: redacted message + session A->>T: get_order(email) T->>A: order 4412, delivered late A->>R: search("late delivery refund policy") R->>A: policy chunks + citations alt refund within auto-approve limit A->>T: create_refund(4412, 30.00) T->>A: refund_id A->>U: refund issued + policy link else above limit or low confidence A->>H: handoff with summary + draft reply H->>U: approves / edits and sends end

Typical support-agent run — look up the record, ground the answer in policy, act within limits, and hand off with a draft when the action is risky

Build path

Ship one narrow agent to production first; add channels and tools only after the eval gate is green

Five agent patterns (and when to use each)

PatternTypical use caseCore piecesComplexity
1. Knowledge agent (RAG)Internal knowledge base, policy Q&A, product docs assistantIngestion, hybrid search, reranker, citations, ACL filteringLow–medium
2. Action agent (tool calling)AI customer support agent, order status, refunds, CRM updatesRAG + typed tools via MCP, approval thresholds, audit logMedium
3. Voice agentAI receptionist, appointment booking, lead qualification callsTwilio/SIP, streaming STT/TTS (Vapi, Retell, LiveKit), short tool calls, SMS follow-upMedium
4. Workflow agentInvoice/document processing, email triage, lead enrichmentn8n or Temporal workflow with LLM steps, structured outputs, retriesLow–medium
5. Multi-agent systemResearch + drafting pipelines, complex ops spanning many systemsSupervisor + specialist agents (LangGraph), shared state, per-agent permissionsHigh

Rule of thumb: start at the lowest pattern that solves the problem. Most “we need a multi-agent system” requests are really pattern 2 with good retrieval. Move to pattern 5 only when one agent’s tool list, permissions, or prompt become unmanageable.

Indicative mix of inbound AI agent requests by pattern (stackcone pipeline, 2026)

Recommendation: Python (FastAPI) agent service on LangGraph or a vendor agent SDK (Claude Agent SDK / OpenAI Agents SDK), Postgres + pgvector for state and retrieval, integrations exposed as MCP servers, a model gateway for provider switching, and Langfuse for tracing and evals. For SMB workflow agents, n8n (self-hosted in the EU if needed) keeps the automation visible to the client.

LayerTechnologyWhy
Agent runtimeLangGraph, Claude Agent SDK, or OpenAI Agents SDKExplicit state graph, retries, human-interrupt nodes; SDKs are faster for single-agent builds
ModelsClaude, GPT, Gemini; Llama/Mistral/Qwen for self-hostedRoute by task: small fast model for classification, frontier model for planning
Model gatewayLiteLLM or cloud-native (Bedrock, Azure OpenAI, Vertex)One API, regional endpoints, fallback, per-tenant cost caps
Retrievalpgvector or OpenSearch + BM25 hybrid, Cohere/Voyage rerankerHybrid + rerank beats pure vector search on business documents
IntegrationsMCP servers (TypeScript or Python)Reusable across agents, IDEs and chat clients; clear per-tool scopes
VoiceVapi or Retell (fast), LiveKit Agents (custom), Twilio numbersSub-second turn-taking and barge-in are hard to build from scratch
Workflowsn8n (SMB), Temporal (engineering teams)Durable retries and schedules outside the LLM loop
Observability & evalsLangfuse or LangSmith, plus a pytest eval suite in CITrace every run; block releases that regress accuracy
FrontendNext.js widget + admin console; Slack/Teams appsStreaming responses (SSE), feedback buttons, source previews
HostingAWS / GCP / Azure in the client’s account and chosen regionData residency and procurement are simpler when it runs in their cloud

Why MCP for integrations? A HubSpot or Zendesk MCP server built once can be used by the support agent, the sales agent, and staff inside Claude or ChatGPT. It also forces clean tool contracts — name, typed arguments, scope — which is exactly what security reviews ask for.

Why not start with a no-code agent builder? They are fine for prototypes. Production buyers usually need custom retrieval, permission-aware data, eval gates and hosting control, which is where a custom build pays off. A common path is prototype in n8n, then move the core loop to code.

Typical monthly run-cost split for a mid-volume support agent (illustrative)

Component design

1 — Knowledge layer (RAG)

2 — Action layer (tools + MCP)

3 — Agent core

4 — Channels

ChannelAdapterNotes
Web chatNext.js widget over SSEStreaming, source cards, thumbs up/down
Slack / TeamsBot app + events APIUses the employee’s identity for permission-aware answers
WhatsApp / SMSTwilio or WhatsApp Cloud APITemplate messages for outbound; opt-in tracking
VoiceVapi / Retell / LiveKit on Twilio numbersShort tool calls, SMS follow-up for links, recording disclosure
EmailInbound parse webhookDraft-for-approval by default

5 — Trust layer

GDPR, UK GDPR & EU AI Act

European and UK buyers search for “GDPR compliant AI chatbot” and “EU AI Act compliant AI” because procurement will ask. The architecture above supports this by design:

RequirementHow the design handles it
Data residencyIn-region model endpoints (Bedrock, Azure OpenAI or Vertex AI in the required region) or self-hosted open-weight models; vector DB and logs in the same region
No training on customer dataEnterprise/API terms with zero data retention where offered; documented in the DPA
Data minimisationPII redacted before tracing; configurable log retention (e.g. 30 days)
Right of access / erasureConversations keyed by user ID; delete endpoint cascades to logs and memory
Transparency (AI Act)Clear “you are talking to an AI” disclosure in chat and at the start of voice calls
Human oversightApproval queue and handoff for consequential decisions; no fully automated decisions with legal effect
Record keepingTrace store with model version, prompt version, tools called, and reviewer actions
Calls & messagingRecording disclosure, PECR/TCPA-aware outbound rules, opt-out handling for SMS/WhatsApp

This is engineering guidance, not legal advice — the client’s DPO or counsel should confirm risk classification for their specific use case.

Implementation plan

Phase 1 — Discovery & golden set (week 1)

Pick one use case with clear volume and ROI (e.g. top 20 support intents). Collect 100–300 real questions and expected answers/actions. Map systems, APIs, permissions, and data residency needs. Agree success metrics: resolution rate, accuracy, CSAT, cost per conversation.

Risk: No API access to key systems — confirm credentials and sandbox accounts in week 1. Rollback: none needed; discovery output stands alone as a spec.

Phase 2 — Knowledge layer (week 2)

Ingest sources with access metadata, build hybrid search + reranker, and measure recall on the golden set. Fix chunking and missing content before touching prompts.

Risk: Stale or contradictory documents — surface conflicts to content owners. Rollback: read-only; no production impact.

Phase 3 — Tools & agent core (weeks 2–3)

Build 2–4 MCP tools (read first, then one write tool with thresholds). Wire the agent loop, state, and model gateway. Add scenario tests for every tool path.

Risk: Over-broad API scopes — create dedicated service accounts. Rollback: write tools behind a feature flag.

Phase 4 — Channel, guardrails & eval gate (weeks 3–4)

Ship the first channel (usually web widget or helpdesk sidebar), PII redaction, injection filtering, handoff flow, and CI eval suite. Release only when accuracy and tool-call tests pass agreed thresholds.

Risk: Accuracy plateau — expand golden set and fix retrieval before prompt tuning. Rollback: widget toggle; traffic back to humans.

Phase 5 — Shadow mode & go-live (weeks 4–5)

Agent drafts replies that humans approve for 1–2 weeks; compare against human answers. Then turn on auto-resolve for intents that clear the bar, keeping approval for the rest.

Risk: Edge cases not in golden set — feed every human correction back into evals. Rollback: per-intent switch back to draft-only.

Phase 6 — Expand channels & hand over (week 6+)

Add voice, Slack/Teams or WhatsApp on the same core; add more tools. Deliver runbook, dashboards, prompt/version registry, and admin training.

Risk: Cost growth with volume — per-tenant caps and model routing. Rollback: channels are independent adapters; disable one without affecting others.

Evals, monitoring & ops

SignalSourceCadence
Resolution rate (no human needed)Conversation outcomesDaily dashboard; weekly review
Answer faithfulness & citation rateLLM-judge + sampled human reviewEvery release (CI) + weekly sample
Tool-call correctnessScenario test suiteEvery release
Handoff rate by intentTrace storeWeekly — target intents above 30% for improvement
Latency p50 / p95TracingReal-time alerts
Cost per conversationGateway token + tool usageWeekly; alert on 30% jumps
CSAT / thumbs feedbackWidget + helpdeskWeekly
Knowledge freshnessIngestion job statusDaily; alert on failed syncs

More on this in AI agent testing and evaluation strategies.

Indicative auto-resolution ramp — shadow mode first, then per-intent enablement (illustrative)

Cost & effort estimate

ScopeTimelineIndicative cost (USD; GBP/EUR similar)
Pilot: 1 use case, 1 channel, RAG + 2–4 tools, eval gate3–6 weeks8k–25k
AI receptionist / voice agent with booking + CRM3–5 weeks10k–30k
Production multi-channel agent with SSO, admin console, EU hosting8–14 weeks25k–80k
Multi-agent system across several departments12+ weeks60k+
Run costs (models, hosting, voice minutes, tracing)monthly300–3,000 depending on volume
Tuning retainer (evals, prompts, new intents)monthly1k–4k

Ranges depend on API quality of existing systems, number of integrations, compliance scope, and how clean the source documents are.

FAQ

How much does it cost to build a custom AI agent for a business?

A focused pilot usually costs 8k–25k over 3–6 weeks. Multi-channel production agents with voice, evals and SSO land in the 25k–80k range, plus monthly run and tuning costs.

What is the difference between an AI chatbot and an AI agent?

A chatbot answers. An agent also acts — it calls your CRM, helpdesk, calendar or billing APIs, follows multi-step plans, and hands off to a human when unsure.

Do I need a multi-agent system?

Usually not at first. One well-scoped agent with good retrieval and a few tools covers most use cases. Split into multiple agents when permissions or tool count make one agent unmanageable.

Can an AI agent be GDPR and EU AI Act compliant?

Yes — EU-region or self-hosted models, zero data retention, PII redaction, a DPA, AI disclosure, and human oversight for consequential decisions.

Who builds custom AI agents like this?

stackcone is an AI agent development company founded by Amar Kumar. We build RAG chatbots, AI customer support agents, AI voice agents and MCP integrations for businesses, with fixed-scope pilots and full handover. See our portfolio or contact us.

What is MCP and why does it matter?

The Model Context Protocol is an open standard for exposing tools and data to AI models. MCP-based integrations work across Claude, ChatGPT, IDEs and custom agents, so you are not locked into one framework. See AI agents and MCP in practice.

Glossary

TermMeaning
AI agentAn LLM-driven program that plans, calls tools, and acts toward a goal, not just answers text
Agentic AIUmbrella term for systems where models take multi-step actions with some autonomy
RAGRetrieval-augmented generation — answering from retrieved company documents with citations
Tool callingThe model returns a structured request to run a function (e.g. create_refund); your code executes it
MCPModel Context Protocol — open standard for connecting models to tools and data sources
Human-in-the-loopA person approves or edits agent actions above a risk threshold
Golden setCurated real questions and expected answers/actions used to measure accuracy
Shadow modeAgent drafts responses that humans review before anything reaches the customer
Model gatewayA single API in front of several LLM providers for routing, fallback, and cost control