Section 191 · Appendices, Tools, Templates, and Reference
Appendix: Testing SKILL.md
A SKILL.md file is not just documentation. It is executable intent for an AI coding agent, so it needs to be tested like product behavior.
SKILL.mdskill triggeragent instructionsprogressive disclosureskill evaluation
What to do
- Start by testing discoverability.
- Test instruction clarity.
- Compare agent runs with and without the skill.
- Keep a small suite of realistic tasks that should trigger the skill.
- Run them when the skill changes.
Evidence to preserve
- Track skill version, trigger terms, tool dependencies, success criteria, conflicting instructions, and replay results.
- Preserve the inputs, versions, configurations, raw outcomes, and results for SKILL.md, skill trigger, agent instructions, progressive disclosure needed to reproduce work on Appendix: Testing SKILL.md.
- Report results for SKILL.md, skill trigger, agent instructions, progressive disclosure by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
When the system matters, treat SKILL.md as versioned agent behavior. Track skill version, trigger terms, tool dependencies, success criteria, conflicting instructions, and replay results. A skill that cannot be evaluated is just a wish written in Markdown.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Appendix: Testing SKILL.md." Testing AI Knowledge Edition, section 191.
https://jarbon.ai/testing-ai/knowledge/ch191-skill-md.html