A support agent demo takes an afternoon: send the customer’s message to a language model with a good prompt and post the answer back. A support agent that can run against real tickets, with real orders, real money and customers who are already upset, looks very different.
This article walks through that shape, step by step, in Python. It’s based on AI Central, the customer-support platform I designed for a large online retailer. I built its proof of concept and then led the developers who took it to production. It answers tickets such as “where is my order?”, delivery delays, damaged packages and technical problems by reading live order data and following a curated knowledge base. Since April 2026 it has processed more than 61,000 tickets and sent over 110,000 automated replies.
The code below is simplified and rewritten for this article, but the architecture, the numbers and the lessons come from the real system.
The shape of the system
The most important decision is what the model is not allowed to do. In AI Central the model doesn’t decide which systems to call, doesn’t loop until it feels done and doesn’t send anything on its own. It works inside a deterministic workflow:
helpdesk webhook ──▶ FastAPI ──▶ SQS ──▶ worker
│
┌────────────────────┘
▼
security pre-check ─▶ collect order data (ERP)
│
▼
classify ─▶ retrieve policy (pgvector) ─▶ generate
│
▼
reply · ask the customer · hand off to a human
Every arrow is ordinary code. The model is called in a few well-defined places: to classify, to check for abuse and to write the reply. That makes the system testable, auditable and much cheaper to debug.
Step 1: accept webhooks fast, process them later
Helpdesk tools deliver events through webhooks and retry when they don’t get a quick answer. Processing a ticket can take many seconds (ERP calls, retrieval, one or more model calls), so the webhook handler should only validate, enqueue and return 202 Accepted.
from fastapi import FastAPI, Depends, status
from pydantic import BaseModel
app = FastAPI()
class TicketEvent(BaseModel):
ticket_id: int
group_id: int
kind: str # "new" or "customer_reply"
@app.post("/webhooks/tickets", status_code=status.HTTP_202_ACCEPTED)
async def receive(event: TicketEvent, _: None = Depends(verify_token)):
await queue.send(event.model_dump())
return {"queued": True}
Two details mattered in practice. First, authenticate the webhook with a token kept in configuration, never in code. Second, be forgiving with the payload: helpdesk automations aren’t always good at producing valid JSON, and AI Central’s parser accepts JSON, form data and query strings for that reason.
Step 2: a worker with retries and a dead-letter queue
The worker long-polls the queue and processes several tickets concurrently. AI Central uses SQS with a 20-second long poll, a 900-second visibility timeout and a default concurrency of four.
import asyncio
semaphore = asyncio.Semaphore(4)
async def worker_loop():
while True:
messages = await queue.receive(max_messages=4, wait_seconds=20)
await asyncio.gather(*(handle(m) for m in messages))
async def handle(message):
async with semaphore:
try:
await process_ticket(message.body)
await queue.delete(message)
except IntegrationError:
await retry_later(message) # transient: ERP or helpdesk failed
except Exception:
await send_to_dead_letter(message)
The retry policy is deliberately narrow. Only integration failures, such as a timeout from the ERP or a 5xx from the helpdesk, are retried: up to three times, with a delay of 2 ** n seconds capped at 900. Anything else goes to a dead-letter queue and a failed_tasks table, because retrying a bug three times only produces three failures.
Rate limits are their own category. The helpdesk API allows a limited number of calls per minute, so AI Central keeps a sliding-window limiter below that quota and sends throttled tickets to the dead-letter queue tagged rate_limit. A small consumer puts them back after the Retry-After period, with some random jitter so they don’t all return at once.
Step 3: one ticket at a time, and never the same reply twice
Customers often send two messages in a row. Without protection, two workers can process the same ticket simultaneously and both reply.
AI Central guards against this at two levels:
_processing: set[int] = set()
_lock = asyncio.Lock()
async def run_exclusively(ticket_id: int, fn) -> bool:
async with _lock:
if ticket_id in _processing:
return False
_processing.add(ticket_id)
try:
await fn()
return True
finally:
async with _lock:
_processing.discard(ticket_id)
That per-ticket lock lives in the worker process. It’s enough for a single worker. With several worker processes, the same idea needs a shared lock in Redis or in the database, and that’s one of the first improvements I’d make.
The second guard sits at the very end, right before sending: an idempotency key derived from the reply text.
import hashlib
def should_send(redis, ticket_id: int, text: str) -> bool:
normalized = " ".join(text.split())
digest = hashlib.sha256(normalized.encode()).hexdigest()[:16]
return bool(redis.set(f"ticket_sent:{ticket_id}:{digest}", "1", nx=True, ex=120))
If the same reply was sent to the same ticket in the last two minutes, it isn’t sent again. The agent also checks that the last public message came from the customer, so it never talks over a human agent who replied in the meantime.
Step 4: model the conversation as a state machine
This is the decision that made the system manageable. Instead of an open-ended agent loop, each ticket has an explicit state, stored in the database:
from enum import StrEnum
class TicketStep(StrEnum):
NEW_TICKET = "new_ticket"
COLLECTING_DATA = "collecting_data"
AWAITING_ORDER_NUMBER = "awaiting_order_number"
AWAITING_DAMAGE_PHOTOS = "awaiting_damage_photos"
AWAITING_NEW_ADDRESS = "awaiting_new_address"
WAITING_CUSTOMER = "waiting_customer"
COMPLETED = "completed"
REDIRECTED = "redirected" # handed off to a human
In AI Central most states are awaiting_* states, each one a specific question the agent asked and is waiting on. The rest cover new tickets, data collection, reply generation, waiting for the customer and the two ways a ticket ends.
New messages go through a linear pipeline of phases. Each phase either lets the ticket continue or stops the run by returning the next state:
from collections.abc import Awaitable, Callable
Phase = Callable[[TicketContext], Awaitable[TicketStep | None]]
PIPELINE: list[Phase] = [
security_precheck,
extract_order_identifiers,
load_order_from_erp,
match_requester_to_order,
categorize,
retrieve_policy,
handle_special_flows, # e.g. lost shipment, partial delivery
generate_reply,
]
async def run_pipeline(ctx: TicketContext) -> TicketStep:
for phase in PIPELINE:
next_step = await phase(ctx)
await record_event(ctx, phase.__name__, next_step)
if next_step is not None:
return next_step
return TicketStep.WAITING_CUSTOMER
AI Central builds this pipeline with LangGraph: 18 nodes connected by conditional edges that stop when a phase returns a state. The idea doesn’t depend on a framework, though. A list of async functions works.
Replies to an awaiting_* state skip the pipeline and go straight to the handler for that question. If the agent asked for photos of a damaged package, the next message is checked for photos first.
Every phase also appends a row to a ticket_flow_events table: phase, state before and after, and a sanitized payload (at most 35 keys, strings truncated to 400 characters). The current state lives in a regular row; the events are an append-only timeline. When someone asks “why did the agent say that?”, the answer is a query away.
Step 5: guardrails before the model sees anything
Support messages are untrusted input. Before any generation, AI Central runs two checks.
The first is plain code: truncate the message (5,000 characters), strip role markers and delimiters that look like prompt syntax, and flag known injection patterns.
The second is a small, cheap model call with structured output. Using instructor and a Pydantic model, the classifier can only return one of the allowed values:
from typing import Literal
import instructor
from openai import AsyncOpenAI
from pydantic import BaseModel
client = instructor.from_openai(AsyncOpenAI())
class MessageCheck(BaseModel):
verdict: Literal["ok", "attack", "serious_complaint"]
reason: str
async def precheck(message: str) -> MessageCheck:
return await client.chat.completions.create(
model="gpt-4o-mini",
temperature=0,
response_model=MessageCheck,
messages=[
{"role": "system", "content": PRECHECK_PROMPT},
{"role": "user", "content": message[:4000]},
],
)
An attack closes the ticket with a security tag. A serious complaint, such as a threat to go to the consumer protection agency, a lawyer or a lawsuit, goes straight to a human. No model should be negotiating with an upset customer who mentions a lawyer.
Step 6: get the facts with code, not with tool calling
Many agent tutorials give the model tools and let it decide when to look up an order. AI Central doesn’t. Order and shipment lookups are fixed pipeline phases: extract the order number or invoice number from the message, call the ERP (with a 30-second timeout and retries on 5xx), and turn the result into a compact text summary for the prompt.
This has three advantages. The lookups always happen, and always the same way. They’re testable without a model. And the model never sees a tool it could call with the wrong arguments.
It also lets you control exactly what the model sees. One of AI Central’s architecture decisions removed “ready for pickup” tracking states from the context entirely, because the model kept inventing pickup addresses when it saw them. Sometimes the best guardrail is not showing the data.
Step 7: retrieval with pgvector
Policies change: what to do when a delivery is late, when a package arrives damaged, when an address is wrong. AI Central keeps them in a markdown knowledge base, split into sections, each with a title (one of the ticket categories) and a line of keywords.
Each section is embedded with text-embedding-3-small (1,536 dimensions) and stored in PostgreSQL with pgvector:
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE knowledge (
id bigserial PRIMARY KEY,
section_title text NOT NULL UNIQUE,
content text NOT NULL,
embedding vector(1536) NOT NULL
);
Search uses cosine distance:
SEARCH = """
SELECT section_title, content, 1 - (embedding <=> :query) AS similarity
FROM knowledge
WHERE (:category IS NULL OR lower(section_title) = lower(:category))
AND 1 - (embedding <=> :query) >= :min_similarity
ORDER BY embedding <=> :query
LIMIT :limit
"""
The interesting part is the two-pass retrieval:
async def retrieve(query_vec, category):
first = await search(query_vec, category=category, limit=2, min_similarity=0.6)
if first and first[0].similarity >= 0.70:
return first[0], category
fallback = await search(query_vec, category=None, limit=2, min_similarity=0.6)
if fallback and fallback[0].section_title in VALID_CATEGORIES:
# Retrieval disagrees with the classifier: trust the knowledge base.
return fallback[0], fallback[0].section_title
return (fallback[0] if fallback else None), category
The first pass searches only the section for the category the classifier chose. If the best match is weak (below 0.70), the second pass searches the whole knowledge base. If it finds a strong section for a different category, the ticket is reclassified. Retrieval becomes a second opinion on routing.
The query text is built from the last customer message (up to 500 characters), the subject and the start of the description; on follow-ups, only the last message.
A note on scale: the knowledge base is small, so AI Central doesn’t use an approximate index (HNSW or IVFFlat). A sequential scan over a few dozen vectors is instant. Add an index when the table grows, not before.
This replaced an earlier attempt at fine-tuning. For about a day the project had a fine-tuning pipeline with a training set of 300 examples. Moving to retrieval meant policies could be updated by editing markdown, every answer could be traced to a section, and nothing had to be retrained.
Step 8: let the model write, but not decide
The reply is generated with the order summary, the retrieved policy and the recent conversation in the prompt. AI Central uses gpt-4o-mini as the primary model.
Sometimes the reply implies an action: start a lost-shipment flow, ask for a new address, hand off to a person. The model can signal that only through a small, fixed block at the end of its output, which the code strips and validates against an allowlist:
VALID_WORKFLOWS = {
"lost_shipment",
"address_change",
"damaged_package",
"human_handoff",
# ...
}
def extract_workflow(text: str) -> tuple[str, str | None]:
reply_lines, workflow = [], None
for line in text.splitlines():
if line.lower().startswith("workflow:"):
slug = line.split(":", 1)[1].strip().strip("'\"").lower()
if slug in VALID_WORKFLOWS:
workflow = slug
continue
reply_lines.append(line)
return "\n".join(reply_lines).strip(), workflow
Anything outside the allowlist is ignored, and the prompt tells the model to omit the block when in doubt. A special marker in the output means “escalate to a human now”, and it always wins.
Step 9: keep cost and availability under control
Two techniques made the biggest difference.
A cache of reply plans, not replies. Caching literal answers doesn’t work for support, because every customer, order and date is different. AI Central caches blueprints instead: a reusable, PII-free plan for a type of reply (intent, goal, up to eight steps, dynamic fields, special rules).
SELECT blueprint, 1 - (embedding <=> :query) AS similarity
FROM response_blueprints
WHERE lower(category) = lower(:category)
AND playbook_version = :version
AND 1 - (embedding <=> :query) >= 0.86
ORDER BY embedding <=> :query
LIMIT 1;
On a hit, the prompt is rendered in a compact form: the blueprint replaces the raw policy text, the conversation history shrinks from five messages to two, and the output budget drops from 1,000 to 600 tokens. On a miss, a new blueprint is generated and stored. Bumping playbook_version invalidates every plan when policies change. At the time of writing, the cache holds 2,298 plans and has served 1,985 hits.
A fallback provider. Model APIs fail. AI Central retries the primary provider three times with exponential backoff and then falls back to Anthropic’s Claude:
async def complete(messages, **kwargs) -> str:
last_error = None
for attempt in range(3):
try:
response = await openai.chat.completions.create(messages=messages, **kwargs)
return response.choices[0].message.content
except Exception as error:
last_error = error
await asyncio.sleep(2 ** attempt)
return await anthropic_fallback(messages, reason=str(last_error), **kwargs)
One subtlety: if the output fails the content checks, that’s not an availability problem, so it raises an error instead of falling back. Asking a second model to try again doesn’t fix a suspicious answer.
Model errors get their own handling too. A timeout, a refused request and an unparseable answer each lead to a defined outcome (retry, fallback or handoff) instead of a generic exception that could leave a customer without an answer.
Every call logs its token usage and estimated cost, and every trace in LangSmith carries the ticket id. When the bill or the behavior changes, you can find out why.
Step 10: know when to stop
An agent that never gives up is worse than no agent. AI Central has explicit limits:
- A small, fixed number of automated replies per ticket.
- A maximum number of attempts for each question the agent asks before handing off.
- Immediate handoff on serious complaints, on ambiguity about what the customer chose, or when an order exists but has no shipment.
- Idle tickets are resolved automatically after a few days.
Handing off well is a feature. The human agent receives the ticket with the collected data and the full timeline of what the agent did.
From one agent to a platform
AI Central started with a single support area. When a second one arrived (technical support, with a very different conversation), copying the pipeline would have produced two systems drifting apart. Instead, the codebase was split into a kernel and modules.
The kernel owns everything that isn’t domain logic: webhooks, the queue, retries, external clients, persistence and locking. Each module owns its workflow, prompts and knowledge base, behind a small contract:
from typing import Protocol
class ModuleProcessor(Protocol):
name: str
async def process_new(self, ctx: TicketContext) -> None:
"""First run after the ticket is stored."""
async def process_reply(self, ctx: TicketContext) -> None:
"""A customer reply on an existing ticket."""
MODULE_PROCESSORS: dict[str, ModuleProcessor] = {
"delivery": DeliveryProcessor(),
"tech_support": TechSupportProcessor(),
}
A new ticket is routed to a module once, and the module’s name is stored with it, so later replies go back to the same workflow even if the ticket moves between teams. The modules don’t have to look alike: the delivery module is a LangGraph pipeline, while the technical-support module is a smaller hand-written workflow with one handler per step. The only rule is that no module reaches into another one.
Testing and evaluation
The deterministic parts are tested like any other code. AI Central has around 250 unit tests covering routing-block parsing, state transitions, prompt gating, extractors and ERP validators. None of them call a model.
For the model-dependent parts, the most useful practice came from a sibling project, a support assistant for internal ERP questions. It uses a golden set of real questions with expected answers, including cases where the correct behavior is to not answer. Measuring retrieval recall separately from correct abstention, and checking faithfulness with an LLM judge, catches the failure that matters most in support: a confident answer that isn’t in the policy.
Lessons
- Make the model a component, not the controller. A state machine and a pipeline of plain functions are easier to reason about, test and audit than an open-ended loop.
- Fetch facts with code. Deterministic lookups beat tool calling when the set of lookups is known.
- Control what the model sees. Removing misleading data from the context fixed a hallucination that prompting didn’t.
- Use retrieval as a second opinion. A two-pass search can correct the classifier.
- Cache plans, not answers.
- Design the exits. Reply limits, handoff rules and idempotent delivery are what make an agent safe to leave running.
- Shared state needs shared locks. An in-process lock is fine for one worker; scale out, and it has to move to Redis or the database.
The model is the most visible part of a support agent, and probably the smallest. Most of the work, and most of what makes it trustworthy, is ordinary backend engineering around it.