Section 074 · Chapter 10, Anti-Patterns That Create False Confidence
Anti-Patterns: The Golden Answer Problem
Many AI tasks do not have one correct answer, and pretending they do creates bad evals.
golden answergolden answer problem
What to do
- Use multiple reference answers, required-fact extraction, rubric scoring, pairwise preference, and human calibration.
- Treat exact-match accuracy as one tool, not the default metric for open-ended tasks.
- Define runnable checks that exercise golden answer and golden answer problem.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for golden answer, golden answer problem needed to reproduce work on Anti-Patterns: The Golden Answer Problem.
- Report results for golden answer, golden answer problem by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Use multiple reference answers, required-fact extraction, rubric scoring, pairwise preference, and human calibration. Treat exact-match accuracy as one tool, not the default metric for open-ended tasks.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Anti-Patterns: The Golden Answer Problem." Testing AI Knowledge Edition, section 74.
https://jarbon.ai/testing-ai/knowledge/ch074-golden-answer-problem.html