challenge is agent-read markdown (skill) from pedrohcgs/claude-code-my-workflow: Stress-test a finding against the choices you did not make. Enumerates the discrete forks a competent analyst could have taken (measure definition, sample filter, control set, clustering level, weighting, functional form), runs the specification grid, and reports the distribution rather than a point estimate — then attacks the identifying assumption with named, computable sensitivity statistics. Use when the user says "is this robust", "challenge this result", "specification curve", "multiverse".
Indexed from public GitHub and served as immutable, content-addressed versions. Install it pinned to an exact SHA-256 with the mdr CLI, and every file is verified before it reaches your agent: the main file against the SHA-256 recorded here, the others against the git hashes of its source commit. The deterministic audit below grades the latest version, and the same checks always give the same file the same grade.
What the file says
# Challenge — does the result survive the choices you didn't make?
A single specification is one draw from a distribution you never looked at.
**Why this exists, measured rather than asserted.** In a controlled study, 150 autonomous
agents were given the same data and the same questions. Effect-size interquartile ranges
reached **~10.7 %/yr**, and the spread concentrated in **discrete measure-choice forks** — not
in estimation noise. *Within* a measure family, agents agreed to ~0.25 %/yr. Two findings from
that study shape this skill:
- **AI peer review left the spread essentially unchanged.** Review catches errors; it does
**not** reduce analytical-choice variance. A clean referee report is not robustness.
- Exposure to exemplar papers collapsed the spread by 80–99 % — **convergence by imitation, not
by correctness.** Herding is not agreement.
So the spread has to be *measured*, not reviewed away.
## Preconditions
- A working baseline specification that runs and produces the headline estimate.
- The estimate's **estimand stated in words** — "the ATT for units treated in 2015, over
…
Pin to a label to follow the author's releases, or to a sha256 for exact bytes. Either way the resolved hash is written to mdr.lock, and mdr install fetches those bytes again and checks them, so it installs them exactly or fails.
GET https://markdownregistry.com/api/v1/artifacts/art_lzs2h3ee2fqhgqrt
GET https://markdownregistry.com/api/v1/resolve?ref=pedrohcgs/claude-code-my-workflow/challenge
GET https://markdownregistry.com/api/v1/blob/aac4117e29ad36394bc7915c30f2ef7a0147a0eff6417c979bfc7739ae113c1b
Your agent does the legwork. You hear about the deals worth your word. Hand yours the standing instructions at modelranch.com and it joins the network that reads files like this one.
pedrohcgs/claude-code-my-workflow · .claude/skills/adjudicate-review/SKILL.md · Turn an incoming set of findings — from an AI reviewer, a referee report, a code review, a linter, or a second model —…
pedrohcgs/claude-code-my-workflow · .claude/skills/audit-reproducibility/SKILL.md · Enforce the replication-protocol.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata…
pedrohcgs/claude-code-my-workflow · .claude/skills/blast-radius/SKILL.md · Before and after changing anything shared — a function's return value, a signature, a schema, a label set, a config…
pedrohcgs/claude-code-my-workflow · .claude/skills/capture-environment/SKILL.md · Snapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and…
pedrohcgs/claude-code-my-workflow · .claude/skills/checkpoint/SKILL.md · Save a structured state snapshot before stopping or handing off. Captures the active plan, recent decisions, file…
pedrohcgs/claude-code-my-workflow · .claude/skills/coauthor-brief/SKILL.md · Generate a co-author / collaborator handoff brief for a multi-author, multi-machine project — summarizing what changed…
pedrohcgs/claude-code-my-workflow · .claude/skills/commit/SKILL.md · Stage, commit, push, open a PR, and merge to main. Use ONLY on explicit commit intent — user says "commit", "ship it"…
pedrohcgs/claude-code-my-workflow · .claude/skills/compile-latex/SKILL.md · Compile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex). Use when user says "compile", "build the slides"…
pedrohcgs/claude-code-my-workflow · .claude/skills/compress-session/SKILL.md · Distill the current conversation into a structured note (decisions made, open questions, file pointers with line…
pedrohcgs/claude-code-my-workflow · .claude/skills/context-status/SKILL.md · Show current context status and session health.
Use to check how much context has been used, whether auto-compact…
pedrohcgs/claude-code-my-workflow · .claude/skills/create-lecture/SKILL.md · Create a new Beamer lecture `.tex` from source papers and materials, with notation consistency checks and the project's…
pedrohcgs/claude-code-my-workflow · .claude/skills/credible-claims/SKILL.md · Research-brief + claim-record discipline for delegated or AI-assisted research work. Use when starting any substantive…
oliver-kriska/claude-elixir-phoenix · plugins/elixir-phoenix/skills/challenge/SKILL.md · Challenge mode reviews - rigorous questioning before approving changes. Use when you want thorough scrutiny of Ecto…
athola/claude-night-market · plugins/gauntlet/skills/challenge/SKILL.md · Presents adaptive codebase challenge questions with multiple-choice and trace exercises. Use when testing contributor…
agentii-ai/agentii-investment-intelligence · plugins/vertical-plugins/scenarios/skills/agentii/challenge/SKILL.md · Adversarial verification of research theses — cross-run/cross-thesis contradiction via the entity index, pre-mortem…
fusengine/agents · plugins/ai-pilot/skills/challenge/SKILL.md · Use before a root-cause, done/verified claim, irreversible action, or 2nd-time fix reaches the owner (APEX or plain…
thoughtbot/rails-consultant · skills/challenge/SKILL.md · Pressure-test an assumption, decision, or inherited constraint — Socratic cross-examination that forces you to defend…
martineserios/thebrana · system/skills/challenge/SKILL.md · Adversarial review — Fable 5 stress-tests reasoning, Gemini checks knowledge. Use before plan or architecture decisions.