Section 180 · Appendices, Tools, Templates, and Reference

Appendix: Using Promptfoo

A lightweight eval tool turns prompt checks from vibes into repeatable tests that can run locally, in CI, and before release.

LLM judgeretrievalPromptfoo

What to do

  1. Do not confuse tool output with truth.
  2. Treat Promptfoo as eval infrastructure, not a substitute for evaluation design and not something this book should document feature by feature.
  3. Version configs, lock datasets, track judge model changes, separate exploratory runs from release gates, and periodically compare automated scores against human raters.

Evidence to preserve

  • Include normal user tasks, edge cases, policy boundaries, prior production failures, and adversarial inputs.
  • Preserve the inputs, versions, configurations, raw outcomes, and results for LLM judge, retrieval, Promptfoo needed to reproduce work on Appendix: Using Promptfoo.
  • Report results for LLM judge, retrieval, Promptfoo by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

Treat Promptfoo as eval infrastructure, not a substitute for evaluation design and not something this book should document feature by feature. Version configs, lock datasets, track judge model changes, separate exploratory runs from release gates, and periodically compare automated scores against human raters.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Appendix: Using Promptfoo." Testing AI Knowledge Edition, section 180.

https://jarbon.ai/testing-ai/knowledge/ch180-promptfoo.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400