Section 191 · Appendices, Tools, Templates, and Reference

Appendix: Testing SKILL.md

A SKILL.md file is not just documentation. It is executable intent for an AI coding agent, so it needs to be tested like product behavior.

SKILL.mdskill triggeragent instructionsprogressive disclosureskill evaluation

What to do

  1. Start by testing discoverability.
  2. Test instruction clarity.
  3. Compare agent runs with and without the skill.
  4. Keep a small suite of realistic tasks that should trigger the skill.
  5. Run them when the skill changes.

Evidence to preserve

  • Track skill version, trigger terms, tool dependencies, success criteria, conflicting instructions, and replay results.
  • Preserve the inputs, versions, configurations, raw outcomes, and results for SKILL.md, skill trigger, agent instructions, progressive disclosure needed to reproduce work on Appendix: Testing SKILL.md.
  • Report results for SKILL.md, skill trigger, agent instructions, progressive disclosure by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

When the system matters, treat SKILL.md as versioned agent behavior. Track skill version, trigger terms, tool dependencies, success criteria, conflicting instructions, and replay results. A skill that cannot be evaluated is just a wish written in Markdown.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Appendix: Testing SKILL.md." Testing AI Knowledge Edition, section 191.

https://jarbon.ai/testing-ai/knowledge/ch191-skill-md.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400