Section 165 · Chapter 20, The Practical Playbook
Worked Example: Testing a Customer-Support Chatbot
A full AI quality workflow shows how the pieces of the book fit together.
customer support chatbot
What to do
- Score policy correctness, completeness, groundedness, tone, user actionability, and safety.
- Separate blockers such as privacy leakage, unsupported financial promises, and account-security mistakes.
- Compare the old system, new prompt, new model, and lower-cost model.
- Record model version, prompt version, retrieval snapshot, tool versions, token use, latency, and cost.
- Use an LLM judge, but calibrate it.
Evidence to preserve
- Include production traces, common billing questions, high-risk account recovery cases, prior failures, Spanish-language cases, long angry messages, and adversarial attempts to bypass refund rules.
- Record model version, prompt version, retrieval snapshot, tool versions, token use, latency, and cost.
- Preserve the inputs, versions, configurations, raw outcomes, and results for customer support chatbot needed to reproduce work on Worked Example: Testing a Customer-Support Chatbot.
- Report results for customer support chatbot by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
At scale, the worked example becomes a repeatable release playbook: sample, score, calibrate, slice, cluster, decide, monitor, and feed production failures back into the eval suite.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Worked Example: Testing a Customer-Support Chatbot." Testing AI Knowledge Edition, section 165.
https://jarbon.ai/testing-ai/knowledge/ch165-customer-support-chatbot.html