Support chats resolved by AI and escalated with context
An online outdoor apparel retailer was drowning in order status, returns and billing chats that queued for hours. We built an AI support assistant grounded in their Orders and Payments APIs that answers with live tracking, starts returns and emails labels on its own, and hands billing disputes to a human with the full transcript, order and payment history attached, so nobody repeats themselves.
- 71%
- Chats resolved by AI
- < 20s
- First response
- 4.8
- CSAT on AI chats
What they came with
The same questions arrived every day, and agents answered them by hand from documentation that already carried the answer. The queue was ordered by arrival rather than by difficulty, so a customer with a genuinely hard problem waited behind a run of people asking where a setting lived. Answers varied with whoever picked the ticket up, and there was no reliable way to see which topics were generating the repeat contact.
What the engagement covered
- Grounded answers cited to the Orders API, never guessed
- Real actions taken: returns started, labels emailed, refunds queued
- Escalations arrive with transcript, order and payments already attached
Technical detail
Retrieval over docs and tickets
Documentation and resolved tickets are chunked, embedded and stored in pgvector alongside the application tables, so a retrieved chunk joins back to its source in a single query. Re-embedding runs on publish, so an answer follows the current doc rather than a nightly snapshot.
Citations bound to retrieved chunks
Each answer carries the chunk ids it was generated from, so the citation shown to the customer is the passage the model actually read. If nothing clears the similarity threshold the assistant does not answer, it escalates.
Confidence decides the handover
Escalation is triggered by weak retrieval and low model confidence rather than by keyword matching. The handover writes the conversation summary and the sources tried into the ticket, so the agent starts without asking the customer to repeat themselves.
Every conversation tagged by topic
Conversations are labelled from the retrieved topic and recorded as resolved or escalated, and the per-topic deflection and satisfaction figures are counted from that record. The same tagging shows which documentation to write next.
The stack
Frontend
AI
Data
Runtime
- Practice
- Customer Support Agents
- Sector
- Online Retail
- Shape
- Client engagement
- Stack
- Anthropic, Shopify API, Next.js
Something like this to build?
Tell us what runs today and where it hurts. An engineer reads it and replies.