42 added, 399 removed. Audit A to A.
---
name: writing-skills
description: Use when creating new skills, editing existing skills, or verifying skills work before deployment
---
# Writing Skills
- ## Overview
-
- **Writing skills IS Test-Driven Development applied to process documentation.**
-
- **Personal skills live in your runtime's skills directory** — see [claude-code-tools.md](https://github.com/obra/Superpowers/blob/main/skills/using-superpowers/references/claude-code-tools.md), [codex-tools.md](https://github.com/obra/Superpowers/blob/main/skills/using-superpowers/references/codex-tools.md), [copilot-tools.md](https://github.com/obra/Superpowers/blob/main/skills/using-superpowers/references/copilot-tools.md), or [gemini-tools.md](https://github.com/obra/Superpowers/blob/main/skills/using-superpowers/references/gemini-tools.md) for the path on your runtime. Codex, Copilot CLI, and Gemini CLI all also recognize `~/.agents/skills/` as a cross-runtime alias.
-
- You write test cases (pressure scenarios with subagents), watch them fail (baseline behavior), write the skill (documentation), watch tests pass (agents comply), and refactor (close loopholes).
-
- **Core principle:** If you didn't watch an agent fail without the skill, you don't know if the skill teaches the right thing.
-
- **REQUIRED BACKGROUND:** You MUST understand superpowers:test-driven-development before using this skill. That skill defines the fundamental RED-GREEN-REFACTOR cycle. This skill adapts TDD to documentation.
-
- **Official guidance:** For Anthropic's official skill authoring best practices, see anthropic-best-practices.md. This document provides additional patterns and guidelines that complement the TDD-focused approach in this skill.
-
- ## What is a Skill?
-
- A **skill** is a reference guide for proven techniques, patterns, or tools. Skills help future agents find and apply effective approaches.
-
- **Skills are:** Reusable techniques, patterns, tools, reference guides
-
- **Skills are NOT:** Narratives about how you solved a problem once
-
- ## TDD Mapping for Skills
-
- | TDD Concept | Skill Creation |
- |-------------|----------------|
- | **Test case** | Pressure scenario with subagent |
- | **Production code** | Skill document (SKILL.md) |
- | **Test fails (RED)** | Agent violates rule without skill (baseline) |
- | **Test passes (GREEN)** | Agent complies with skill present |
- | **Refactor** | Close loopholes while maintaining compliance |
- | **Write test first** | Run baseline scenario BEFORE writing skill |
- | **Watch it fail** | Document exact rationalizations agent uses |
- | **Minimal code** | Write skill addressing those specific violations |
- | **Watch it pass** | Verify agent now complies |
- | **Refactor cycle** | Find new rationalizations → plug → re-verify |
-
- The entire skill creation process follows RED-GREEN-REFACTOR.
-
- ## When to Create a Skill
-
- **Create when:**
-
- - Technique wasn't intuitively obvious to you
- - You'd reference this again across projects
- - Pattern applies broadly (not project-specific)
- - Others would benefit
-
- **Don't create for:**
-
- - One-off solutions
- - Standard practices well-documented elsewhere
- - Project-specific conventions (put in your instructions file)
- - Mechanical constraints (if it's enforceable with regex/validation, automate it—save documentation for judgment calls)
-
- ## Skill Types
-
- ### Technique
-
- Concrete method with steps to follow (condition-based-waiting, root-cause-tracing)
-
- ### Pattern
-
- Way of thinking about problems (flatten-with-flags, test-invariants)
-
- ### Reference
-
- API docs, syntax guides, tool documentation (office docs)
-
- ## Directory Structure
-
- ```
- skills/
- skill-name/
- SKILL.md # Main reference (required)
- supporting-file.* # Only if needed
- ```
-
- **Flat namespace** - all skills in one searchable namespace
-
- **Separate files for:**
-
- 1. **Heavy reference** (100+ lines) - API docs, comprehensive syntax
- 2. **Reusable tools** - Scripts, utilities, templates
-
- **Keep inline:**
-
- - Principles and concepts
- - Code patterns (< 50 lines)
- - Everything else
-
- ## SKILL.md Structure
-
- **Frontmatter (YAML):**
-
- - Two required fields: `name` and `description` (see [agentskills.io/specification](https://agentskills.io/specification) for all supported fields)
- - Max 1024 characters total
- - `name`: Use letters, numbers, and hyphens only (no parentheses, special chars)
- - `description`: Third-person, describes ONLY when to use (NOT what it does)
- - Start with "Use when..." to focus on triggering conditions
- - Include specific symptoms, situations, and contexts
- - **NEVER summarize the skill's process or workflow** (see SDO section for why)
- - Keep under 500 characters if possible
-
- ```markdown
- ---
- name: Skill-Name-With-Hyphens
- description: Use when [specific triggering conditions and symptoms]
- ---
-
- # Skill Name
-
- ## Overview
- What is this? Core principle in 1-2 sentences.
-
- ## When to Use
- [Small inline flowchart IF decision non-obvious]
-
- Bullet list with SYMPTOMS and use cases
- When NOT to use
-
- ## Core Pattern (for techniques/patterns)
- Before/after code comparison
-
- ## Quick Reference
- Table or bullets for scanning common operations
-
- ## Implementation
- Inline code for simple patterns
- Link to file for heavy reference or reusable tools
-
- ## Common Mistakes
- What goes wrong + fixes
-
- ## Real-World Impact (optional)
- Concrete results
- ```
-
- ## Skill Discovery Optimization (SDO)
-
- **Critical for discovery:** Future agents need to FIND your skill
-
- ### 1. Rich Description Field
-
- **Purpose:** Your agent reads the description to decide which skills to load for a given task. Make it answer: "Should I read this skill right now?"
-
- **Format:** Start with "Use when..." to focus on triggering conditions
-
- **CRITICAL: Description = When to Use, NOT What the Skill Does**
-
- The description should ONLY describe triggering conditions. Do NOT summarize the skill's process or workflow in the description.
-
- **Why this matters:** Testing revealed that when a description summarizes the skill's workflow, an agent may follow the description instead of reading the full skill content. A description saying "code review between tasks" caused an agent to do ONE review, even though the skill's flowchart clearly showed TWO reviews (spec compliance then code quality).
-
- When the description was changed to just "Use when executing implementation plans with independent tasks" (no workflow summary), the agent correctly read the flowchart and followed the two-stage review process.
-
- **The trap:** Descriptions that summarize workflow create a shortcut agents will take. The skill body becomes documentation agents skip.
-
- ```yaml
- # ❌ BAD: Summarizes workflow - agents may follow this instead of reading skill
- description: Use when executing plans - dispatches subagent per task with code review between tasks
-
- # ❌ BAD: Too much process detail
- description: Use for TDD - write test first, watch it fail, write minimal code, refactor
-
- # ✅ GOOD: Just triggering conditions, no workflow summary
- description: Use when executing implementation plans with independent tasks in the current session
-
- # ✅ GOOD: Triggering conditions only
- description: Use when implementing any feature or bugfix, before writing implementation code
- ```
-
- **Content:**
-
- - Use concrete triggers, symptoms, and situations that signal this skill applies
- - Describe the *problem* (race conditions, inconsistent behavior) not *language-specific symptoms* (setTimeout, sleep)
- - Keep triggers technology-agnostic unless the skill itself is technology-specific
- - If skill is technology-specific, make that explicit in the trigger
- - Write in third person (injected into system prompt)
- - **NEVER summarize the skill's process or workflow**
-
- ```yaml
- # ❌ BAD: Too abstract, vague, doesn't include when to use
- description: For async testing
-
- # ❌ BAD: First person
- description: I can help you with async tests when they're flaky
-
- # ❌ BAD: Mentions technology but skill isn't specific to it
- description: Use when tests use setTimeout/sleep and are flaky
-
- # ✅ GOOD: Starts with "Use when", describes problem, no workflow
- description: Use when tests have race conditions, timing dependencies, or pass/fail inconsistently
-
- # ✅ GOOD: Technology-specific skill with explicit trigger
- description: Use when using React Router and handling authentication redirects
- ```
-
- ### 2. Keyword Coverage
-
- Use words an agent would search for:
-
- - Error messages: "Hook timed out", "ENOTEMPTY", "race condition"
- - Symptoms: "flaky", "hanging", "zombie", "pollution"
- - Synonyms: "timeout/hang/freeze", "cleanup/teardown/afterEach"
- - Tools: Actual commands, library names, file types
-
- ### 3. Descriptive Naming
-
- **Use active voice, verb-first:**
+ Writing a skill IS TDD on process documentation: watch an agent fail the task WITHOUT the skill (RED), write the skill against those exact failures (GREEN), close the loopholes it then invents (REFACTOR). If you never watched it fail without the skill, you don't know the skill teaches the right thing.
- - ✅ `creating-skills` not `skill-creation`
- - ✅ `condition-based-waiting` not `async-test-helpers`
+ Vendored from [obra/Superpowers](https://github.com/obra/Superpowers) (MIT, Jesse Vincent), compressed to house style per `[[vendored-skill-compression]]`.
- ### 4. Token Efficiency (Critical)
+ **REQUIRED BACKGROUND:** superpowers:test-driven-development defines the RED-GREEN-REFACTOR cycle this adapts.
- **Problem:** getting-started and frequently-referenced skills load into EVERY conversation. Every token counts.
+ ## Authoring contract lives in one canonical place
- **Target word counts:**
+ Ordered-by-weight bullets, the behavior-anchored `description` template, description-SDO (WHEN not WHAT), and "match the form to the failure" are the house authoring contract — see `[[skill-authoring-contract]]`, don't restate them. Token-efficient prose style: `[[instruction-compression-playbook]]`. This file covers only what's specific to *creating + testing* a skill.
- - getting-started workflows: <150 words each
- - Frequently-loaded skills: <200 words total
- - Other skills: <500 words (still be concise)
+ ## What a skill is
- **Techniques:**
+ A reference guide for a proven technique, pattern, or tool — reusable, not a narrative of how you solved something once. Three shapes:
- **Move details to tool help:**
+ - **Technique** — a concrete method with steps (`condition-based-waiting`, `root-cause-tracing`).
+ - **Pattern** — a way of thinking about a class of problem (`flatten-with-flags`).
+ - **Reference** — API/syntax/tool docs.
- ```bash
- # ❌ BAD: Document all flags in SKILL.md
- search-conversations supports --text, --both, --after DATE, --before DATE, --limit N
+ ## When to create one
- # ✅ GOOD: Reference --help
- search-conversations supports multiple modes and filters. Run --help for details.
- ```
+ - Create when: the technique wasn't obvious to you, you'd reuse it across projects, and it applies broadly.
+ - Don't create for: one-offs, things well-documented elsewhere, project-specific conventions (put those in the instructions file), or anything a regex/validator can enforce — automate that instead of documenting it.
- **Use cross-references:**
+ ## Iron Law
- ```markdown
- # ❌ BAD: Repeat workflow details
- When searching, dispatch subagent with template...
- [20 lines of repeated instructions]
+ NO SKILL — OR EDIT — SHIPS WITHOUT A FAILING TEST FIRST. Wrote it before testing? Delete it, start from baseline. No exception for "simple additions", "just a section", or "documentation updates". Don't keep untested changes "as reference". Delete means delete.
- # ✅ GOOD: Reference other skill
- Always use subagents (50-100x context savings). REQUIRED: Use [other-skill-name] for workflow.
- ```
+ ## Structure
- **Compress examples:**
+ Two required frontmatter fields: `name` (letters/numbers/hyphens only) and `description` (≤1024 chars total; see [agentskills.io/specification](https://agentskills.io/specification) for the rest). Then a lean body:
```markdown
- # ❌ BAD: Verbose example (42 words)
- your human partner: "How did we handle authentication errors in React Router before?"
- You: I'll search past conversations for React Router authentication patterns.
- [Dispatch subagent with search query: "React Router authentication error handling 401"]
-
- # ✅ GOOD: Minimal example (20 words)
- Partner: "How did we handle auth errors in React Router?"
- You: Searching...
- [Dispatch subagent → synthesis]
- ```
-
- **Eliminate redundancy:**
-
- - Don't repeat what's in cross-referenced skills
- - Don't explain what's obvious from command
- - Don't include multiple examples of same pattern
-
- **Verification:**
-
- ```bash
- wc -w skills/path/SKILL.md
- # getting-started workflows: aim for <150 each
- # Other frequently-loaded: aim for <200 total
- ```
-
- **Name by what you DO or core insight:**
-
- - ✅ `condition-based-waiting` > `async-test-helpers`
- - ✅ `using-skills` not `skill-usage`
- - ✅ `flatten-with-flags` > `data-structure-refactoring`
- - ✅ `root-cause-tracing` > `debugging-techniques`
-
- **Gerunds (-ing) work well for processes:**
-
- - `creating-skills`, `testing-skills`, `debugging-with-logs`
- - Active, describes the action you're taking
-
- ### 5. Cross-Referencing Other Skills
-
- **When writing documentation that references other skills:**
-
- Use skill name only, with explicit requirement markers:
-
- - ✅ Good: `**REQUIRED SUB-SKILL:** Use superpowers:test-driven-development`
- - ✅ Good: `**REQUIRED BACKGROUND:** You MUST understand superpowers:systematic-debugging`
- - ❌ Bad: `See skills/testing/test-driven-development` (unclear if required)
- - ❌ Bad: `@skills/testing/test-driven-development/SKILL.md` (force-loads, burns context)
-
- **Why no @ links:** `@` syntax force-loads files immediately, consuming 200k+ context before you need them.
-
- ## Flowchart Usage
-
- ```dot
- digraph when_flowchart {
- "Need to show information?" [shape=diamond];
- "Decision where I might go wrong?" [shape=diamond];
- "Use markdown" [shape=box];
- "Small inline flowchart" [shape=box];
-
- "Need to show information?" -> "Decision where I might go wrong?" [label="yes"];
- "Decision where I might go wrong?" -> "Small inline flowchart" [label="yes"];
- "Decision where I might go wrong?" -> "Use markdown" [label="no"];
- }
- ```
-
- **Use flowcharts ONLY for:**
-
- - Non-obvious decision points
- - Process loops where you might stop too early
- - "When to use A vs B" decisions
-
- **Never use flowcharts for:**
-
- - Reference material → Tables, lists
- - Code examples → Markdown blocks
- - Linear instructions → Numbered lists
- - Labels without semantic meaning (step1, helper2)
-
- See `graphviz-conventions.dot` in this directory for graphviz style rules.
-
- **Visualizing for your human partner:** Use `render-graphs.js` in this directory to render a skill's flowcharts to SVG:
-
- ```bash
- ./render-graphs.js ../some-skill # Each diagram separately
- ./render-graphs.js ../some-skill --combine # All diagrams in one SVG
- ```
-
- ## Code Examples
-
- **One excellent example beats many mediocre ones**
-
- Choose most relevant language:
-
- - Testing techniques → TypeScript/JavaScript
- - System debugging → Shell/Python
- - Data processing → Python
-
- **Good example:**
-
- - Complete and runnable
- - Well-commented explaining WHY
- - From real scenario
- - Shows pattern clearly
- - Ready to adapt (not generic template)
-
- **Don't:**
-
- - Implement in 5+ languages
- - Create fill-in-the-blank templates
- - Write contrived examples
-
- You're good at porting - one great example is enough.
-
- ## File Organization
-
- ### Self-Contained Skill
-
- ```
- defense-in-depth/
- SKILL.md # Everything inline
- ```
-
- When: All content fits, no heavy reference needed
-
- ### Skill with Reusable Tool
-
- ```
- condition-based-waiting/
- SKILL.md # Overview + patterns
- example.ts # Working helpers to adapt
- ```
-
- When: Tool is reusable code, not just narrative
-
- ### Skill with Heavy Reference
-
- ```
- pptx/
- SKILL.md # Overview + workflows
- pptxgenjs.md # 600 lines API reference
- ooxml.md # 500 lines XML structure
- scripts/ # Executable tools
+ ## Overview # what + core principle, 1-2 sentences
+ ## When to Use # symptoms / use cases; when NOT to; small flowchart only if the decision is non-obvious
+ ## Core Pattern # before/after for techniques & patterns
+ ## Quick Reference # table or bullets for scanning
+ ## Implementation # inline code for simple cases; link a file for heavy reference
+ ## Common Mistakes # what goes wrong + the fix
```
- When: Reference material too large for inline
-
- ## The Iron Law (Same as TDD)
+ SKILL.md is a table of contents (progressive disclosure). Inline only what changes per task; split out when content is heavy:
- ```
- NO SKILL WITHOUT A FAILING TEST FIRST
- ```
+ - **Heavy reference** (100+ lines of API/syntax) → its own file, loaded on demand.
+ - **Reusable tool** (script, template) → its own file.
+ - Keep inline: principles, concepts, code patterns under ~50 lines.
- This applies to NEW skills AND EDITS to existing skills.
+ Cross-reference other skills by name with an explicit marker (`**REQUIRED:** superpowers:test-driven-development`) — never `@path`, which force-loads the file and burns 200k+ context before you need it.
- Write skill before testing? Delete it. Start over.
- Edit skill without testing? Same violation.
+ ## Code examples
- **No exceptions:**
+ One excellent example beats five languages. Make it complete, runnable, from a real scenario, commented with WHY — not a fill-in-the-blank template. You're good at porting; one great example is enough.
- - Not for "simple additions"
- - Not for "just adding a section"
- - Not for "documentation updates"
- - Don't keep untested changes as "reference"
- - Don't "adapt" while running tests
- - Delete means delete
+ ## Flowcharts
- **REQUIRED BACKGROUND:** The superpowers:test-driven-development skill explains why this matters. Same principles apply to documentation.
+ Use a small inline flowchart ONLY for a non-obvious decision point, a process loop where you might stop too early, or an "A vs B" choice. Never for reference material (use tables/lists), code (use markdown blocks), or linear steps (use a numbered list). Style rules: `graphviz-conventions.dot`. Render for a human: `./render-graphs.js ../some-skill [--combine]`.
- ## Testing, Bulletproofing & Checklist
+ ## Testing & bulletproofing
- The full per-skill-type testing methodology, the rationalization-bulletproofing toolkit ("Match the Form to the Failure", loophole-closing, rationalization tables, red-flags), RED-GREEN-REFACTOR-for-skills, anti-patterns, the MANDATORY per-skill deployment checklist, and the discovery workflow live in **[testing-and-bulletproofing.md](testing-and-bulletproofing.md)** — split out to respect this repo's 500-line SKILL.md limit.
+ The full method — pressure scenarios, the RED-GREEN-REFACTOR loop for skills, rationalization tables, red-flags, micro-testing wording against a no-guidance control, and the per-skill deployment gate — lives in **[testing-skills.md](testing-skills.md)**. A worked campaign is in [examples/CLAUDE_MD_TESTING.md](examples/CLAUDE_MD_TESTING.md). Persuasion levers that make a discipline skill bind: persuasion-principles.md.
- Also in this directory: `anthropic-best-practices.md`, `persuasion-principles.md`, `testing-skills-with-subagents.md`, `graphviz-conventions.dot`, `render-graphs.js`, `examples/`.
+ The one-line discipline: discipline/judgment skills get pressure-tested under 3+ stacked pressures; pure reference skills get retrieval-tested instead.
- ## The Bottom Line
+ ## See
- **Creating skills IS TDD for process documentation.** Same Iron Law: no skill without a failing test first. Same cycle: RED (baseline) → GREEN (write skill) → REFACTOR (close loopholes). If you follow TDD for code, follow it for skills.
+ - [testing-skills.md](testing-skills.md) — the testing + bulletproofing method
+ - anthropic-best-practices.md — pointer to Anthropic's public authoring guide + local deltas
+ - persuasion-principles.md — Cialdini levers for compliance under pressure
+ - `[[skill-authoring-contract]]` — the canonical authoring contract (structure, description template, form-to-failure)
+ - `[[instruction-compression-playbook]]` — token-efficient prose style
+ - `[[micro-test-instruction-wording]]` — prove guidance wording binds before shipping it