Project 02
AI Central
A modular AI customer-support platform that answers tickets end to end, grounded in live order data and a curated knowledge base.
- Status
- In production
- Role
- Conceived the platform and built the proof of concept; technical lead of the team that took it to production and keeps extending it.
- Source
- Private codebase
- Helpdesk webhook
- SQS + DLQ
- Module workflow + RAG
- Reply or handoff
Problem
Many support tickets are repetitive questions whose answers depend on order and delivery data, plus policies that change over time. Answering them by hand is slow, and answering them with an unconstrained language model is risky.
Solution
An event-driven Python platform split into a shared kernel and domain modules. The kernel handles webhooks, the queue, external clients, persistence and locking; each module owns its workflow, prompts and knowledge base. It currently runs a delivery-support module built as a LangGraph pipeline and a technical-support module with its own step-by-step workflow. Every ticket ends in a reply, a question back to the customer or a handoff to a person, and every step is recorded for auditing.
Engineering highlights
- In production since April 2026: more than 61,000 tickets processed and 110,000 automated replies sent, with about 85% of tickets receiving at least one answer from the platform.
- Kernel and modules: a small processor contract lets a new support area plug in its own workflow without touching the others.
- Explicit state machines around the model instead of an open-ended agent loop, so every transition can be tested and audited.
- Two-pass retrieval over pgvector: filtered by category first, then across the whole knowledge base, which can also correct the initial classification.
- A semantic cache of reusable reply plans (“blueprints”) instead of literal answers, which keeps prompts short on cache hits. So far it holds 2,298 plans and has served 1,985 hits.
- Structured outputs with Pydantic for every classification, an allowlist for any action the model suggests, and no direct tool calling.
- Layered guardrails: prompt-injection checks, reply limits, idempotent delivery, provider fallback and dedicated handling of model errors, so a failure never reaches the customer.