Section 055 · Chapter 8, Operating AI: Observability, Relevance, and Economics
Prompt and Policy Versioning
Many AI regressions come from changing the instructions around the model, not the model itself.
rubricretrievalpolicy versioningprompt policy versioning
What to do
- Version the system prompt.
- Version policy documents and retrieval indexes.
- Version tools and tool schemas.
- Version judges and rubrics.
- Version data and labels.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for rubric, retrieval, policy versioning, prompt policy versioning needed to reproduce work on Prompt and Policy Versioning.
- Report results for rubric, retrieval, policy versioning, prompt policy versioning by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
At scale, treat prompts, policies, retrieval snapshots, tool contracts, judges, rubrics, datasets, and labels as a single versioned eval bundle. Comparisons across incompatible bundles should be marked as non-equivalent.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Prompt and Policy Versioning." Testing AI Knowledge Edition, section 55.
https://jarbon.ai/testing-ai/knowledge/ch055-prompt-policy-versioning.html