Section 074 · Chapter 10, Anti-Patterns That Create False Confidence

Anti-Patterns: The Golden Answer Problem

Many AI tasks do not have one correct answer, and pretending they do creates bad evals.

golden answergolden answer problem

What to do

  1. Use multiple reference answers, required-fact extraction, rubric scoring, pairwise preference, and human calibration.
  2. Treat exact-match accuracy as one tool, not the default metric for open-ended tasks.
  3. Define runnable checks that exercise golden answer and golden answer problem.

Evidence to preserve

  • Preserve the inputs, versions, configurations, raw outcomes, and results for golden answer, golden answer problem needed to reproduce work on Anti-Patterns: The Golden Answer Problem.
  • Report results for golden answer, golden answer problem by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

Use multiple reference answers, required-fact extraction, rubric scoring, pairwise preference, and human calibration. Treat exact-match accuracy as one tool, not the default metric for open-ended tasks.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Anti-Patterns: The Golden Answer Problem." Testing AI Knowledge Edition, section 74.

https://jarbon.ai/testing-ai/knowledge/ch074-golden-answer-problem.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400