Section 116 · Chapter 15, How Models Work

Useful and Useless LLM Bug Reports

A single bad answer is a clue. It is rarely a complete LLM bug report.

retrievaluseful useless llm bug reports

What to do

  1. Define runnable checks that exercise retrieval and useful useless llm bug reports.
  2. Set acceptable outcomes and blocker failures for retrieval and useful useless llm bug reports before running the evaluation.
  3. Run representative cases for retrieval and useful useless llm bug reports and preserve the failures that would change the decision.

Evidence to preserve

  • Preserve the inputs, versions, configurations, raw outcomes, and results for retrieval, useful useless llm bug reports needed to reproduce work on Useful and Useless LLM Bug Reports.
  • Report results for retrieval, useful useless llm bug reports by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

Convert individual failures into failure classes. A strong LLM bug report names the population, not just the example: "refund escalation hallucination in policy-missing chats" is more useful than "the bot said something wrong."

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Useful and Useless LLM Bug Reports." Testing AI Knowledge Edition, section 116.

https://jarbon.ai/testing-ai/knowledge/ch116-useful-useless-llm-bug-reports.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400