Section 116 · Chapter 15, How Models Work
Useful and Useless LLM Bug Reports
A single bad answer is a clue. It is rarely a complete LLM bug report.
retrievaluseful useless llm bug reports
What to do
- Define runnable checks that exercise retrieval and useful useless llm bug reports.
- Set acceptable outcomes and blocker failures for retrieval and useful useless llm bug reports before running the evaluation.
- Run representative cases for retrieval and useful useless llm bug reports and preserve the failures that would change the decision.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for retrieval, useful useless llm bug reports needed to reproduce work on Useful and Useless LLM Bug Reports.
- Report results for retrieval, useful useless llm bug reports by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Convert individual failures into failure classes. A strong LLM bug report names the population, not just the example: "refund escalation hallucination in policy-missing chats" is more useful than "the bot said something wrong."
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Useful and Useless LLM Bug Reports." Testing AI Knowledge Edition, section 116.
https://jarbon.ai/testing-ai/knowledge/ch116-useful-useless-llm-bug-reports.html