Section 055 · Chapter 8, Operating AI: Observability, Relevance, and Economics

Prompt and Policy Versioning

Many AI regressions come from changing the instructions around the model, not the model itself.

rubricretrievalpolicy versioningprompt policy versioning

What to do

  1. Version the system prompt.
  2. Version policy documents and retrieval indexes.
  3. Version tools and tool schemas.
  4. Version judges and rubrics.
  5. Version data and labels.

Evidence to preserve

  • Preserve the inputs, versions, configurations, raw outcomes, and results for rubric, retrieval, policy versioning, prompt policy versioning needed to reproduce work on Prompt and Policy Versioning.
  • Report results for rubric, retrieval, policy versioning, prompt policy versioning by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

At scale, treat prompts, policies, retrieval snapshots, tool contracts, judges, rubrics, datasets, and labels as a single versioned eval bundle. Comparisons across incompatible bundles should be marked as non-equivalent.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Prompt and Policy Versioning." Testing AI Knowledge Edition, section 55.

https://jarbon.ai/testing-ai/knowledge/ch055-prompt-policy-versioning.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400