Section 180 · Appendices, Tools, Templates, and Reference
Appendix: Using Promptfoo
A lightweight eval tool turns prompt checks from vibes into repeatable tests that can run locally, in CI, and before release.
LLM judgeretrievalPromptfoo
What to do
- Do not confuse tool output with truth.
- Treat Promptfoo as eval infrastructure, not a substitute for evaluation design and not something this book should document feature by feature.
- Version configs, lock datasets, track judge model changes, separate exploratory runs from release gates, and periodically compare automated scores against human raters.
Evidence to preserve
- Include normal user tasks, edge cases, policy boundaries, prior production failures, and adversarial inputs.
- Preserve the inputs, versions, configurations, raw outcomes, and results for LLM judge, retrieval, Promptfoo needed to reproduce work on Appendix: Using Promptfoo.
- Report results for LLM judge, retrieval, Promptfoo by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Treat Promptfoo as eval infrastructure, not a substitute for evaluation design and not something this book should document feature by feature. Version configs, lock datasets, track judge model changes, separate exploratory runs from release gates, and periodically compare automated scores against human raters.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Appendix: Using Promptfoo." Testing AI Knowledge Edition, section 180.
https://jarbon.ai/testing-ai/knowledge/ch180-promptfoo.html