research-review · diff
git:20260710.c3f9813 to git:20260826.363589d
26 added, 1 removed. Audit A to A.
---
name: "research-review"
description: "Get a deep critical review of research from Claude via claude-review MCP. Use when user says \"review my research\", \"help me review\", \"get external review\", or wants critical feedback on research ideas, papers, or experimental results."
---
> Override for Codex users who want **Claude Code**, not a second Codex agent, to act as the reviewer. Install this package **after** `skills/skills-codex/*`.
>
> This reviewer is a different model family from the Codex executor. Every overlay trace/audit records:
>
> ```yaml
> review_independence: cross-family
> acceptance_status: accepted
> ```
# Research Review via `claude-review` MCP (high-rigor review)
> **Claude overlay assurance:** this route is a different model family from the Codex executor and records `review_independence: cross-family` plus `acceptance_status: accepted`.
Get a multi-round critical review of research work from an external LLM with maximum reasoning depth.
## Constants
- **REVIEWER_MODEL = `claude-review`** — Claude reviewer invoked through the local `claude-review` MCP bridge. Set `CLAUDE_REVIEW_MODEL` if you need a specific Claude model override.
- **REVIEWER_BACKEND = `claude-review`** — reviews route through the claude-review MCP (Claude family; cross-family for a Codex executor).
## Context: $ARGUMENTS
## Prerequisites
- Install the base Codex-native skills first: copy `skills/skills-codex/*` into `~/.codex/skills/`.
- Then install this overlay package: copy `skills/skills-codex-claude-review/*` into `~/.codex/skills/` and allow it to overwrite the same skill names.
- Register the local reviewer bridge:
```bash
codex mcp add claude-review -- python3 ~/.codex/mcp-servers/claude-review/server.py
```
- This gives Codex access to `mcp__claude-review__review_start`, `mcp__claude-review__review_reply_start`, and `mcp__claude-review__review_status`.
## Workflow
### Step 1: Gather Research Context
Before calling the external reviewer, compile a comprehensive briefing:
1. Read project narrative documents (e.g., STORY.md, README.md, paper drafts)
2. Read any memory/notes files for key findings and experiment history
3. Identify: core claims, methodology, key results, known weaknesses
### Step 2: Initial Review (Round 1)
Send a detailed prompt with ultra reasoning:
```
mcp__claude-review__review_start:
prompt: |
[Full research context + specific questions]
Please act as a senior ML reviewer (NeurIPS/ICML level). Start from the
assumption that the work is broken somewhere — your job is to find where.
Be adversarial. Trust nothing the author tells you — verify everything
yourself. Identify:
1. Logical gaps or unjustified claims
2. Missing experiments that would strengthen the story
3. Narrative weaknesses
4. Whether the contribution is sufficient for a top venue
- Please be brutally honest.
+
+ === SCOPE LIMITS (these bound what you PROPOSE, never what you look for) ===
+ Report anything that is actually wrong here — including a rare-looking case, if
+ this repo actually produces it. Then keep the fix in scope:
+ 1. This is a RESEARCH-WORKFLOW tool, not a security paper. Verification is
+ welcome; over-defense is not. Assume a cooperating operator on their own
+ machine — a malicious local user is NOT in the threat model.
+ 2. Do NOT propose SHA / hash / content-fingerprint / digest-binding schemes.
+ Reporting a real defect in hashing code that already exists is fine.
+ 3. NO speculative machinery: do not add feature flags, migration frameworks,
+ compat layers, wrappers, pins, or similar mechanisms unless evidence shows
+ a current repo defect they fix or an explicit existing invariant they must
+ preserve. "Load-bearing", "compatibility", and "not scaffolding" are labels,
+ not evidence. Point to the failing path/artifact or invariant, and check the
+ proposal's factual premises, such as whether a named package version exists.
+ 4. NO corner-case obsession: exotic encodings, symlink races, RTL text and
+ millisecond races are out of scope unless you can show the case arises here.
+ 5. Where a rubric or checklist is genuinely needed, do not over-mechanize
+ judgement. A clear sentence a human reads beats a scored table nobody
+ maintains.
+ Exception: code that runs remote commands, starts a network service, or installs
+ an MCP server runs on the user's machine with their credentials — trust-boundary
+ findings there are in scope and the default is strict.
+ Say plainly when something is correct. Do not manufacture findings.
+ Be brutally honest. If, after genuinely trying to break it, the work
+ holds up and is ready, say so clearly.
```
After this start call, immediately save the returned `jobId` and poll `mcp__claude-review__review_status` with a bounded `waitSeconds` until `done=true`. Treat the completed status payload's `response` as the reviewer output, and save the completed `threadId` for any follow-up round.
### Step 3: Iterative Dialogue (Rounds 2-N)
Use `mcp__claude-review__review_reply_start` with the saved completed `threadId`, then poll `mcp__claude-review__review_status` with the returned `jobId` until `done=true` to continue the conversation:
```
mcp__claude-review__review_reply_start:
threadId: [saved reviewer id from Step 2]
prompt: |
Please continue the review using the revised materials below.
Revised files:
- /absolute/path/to/file1
- /absolute/path/to/file2
Focus on unresolved weaknesses and whether the revision actually fixed them.
```
After this start call, immediately save the returned `jobId` and poll `mcp__claude-review__review_status` with a bounded `waitSeconds` until `done=true`. Treat the completed status payload's `response` as the reviewer output, and save the completed `threadId` for any follow-up round.
For each round:
1. **Respond** to criticisms with evidence/counterarguments
2. **Ask targeted follow-ups** on the most actionable points
3. **Request specific deliverables**: experiment designs, paper outlines, claims matrices
Key follow-up patterns:
- "If we reframe X as Y, does that change your assessment?"
- "What's the minimum experiment to satisfy concern Z?"
- "Please design the minimal additional experiment package (highest acceptance lift per GPU week)"
- "Please write a mock NeurIPS/ICML review with scores"
- "Give me a results-to-claims matrix for possible experimental outcomes"
### Step 4: Convergence
Stop iterating when:
- Both sides agree on the core claims and their evidence requirements
- A concrete experiment plan is established
- The narrative structure is settled
### Step 5: Document Everything
Save the full interaction and conclusions to a review document in the project root:
- Round-by-round summary of criticisms and responses
- Final consensus on claims, narrative, and experiments
- Claims matrix (what claims are allowed under each possible outcome)
- Prioritized TODO list with estimated compute costs
- Paper outline if discussed
Update project memory/notes with key review conclusions.
If `— composed: <canonical-report-path>` is explicitly present, fold consensus,
claims matrix, TODOs, and trace links into that report instead of writing a
standalone review document. Without the directive, write the standalone review
as documented; never infer composed mode from an existing file. `— standalone`
always wins. See
[`output-composition.md`](../shared-references/output-composition.md).
### Step 6: Review Tracing
Save a trace for every `mcp__claude-review__review_start`, `mcp__claude-review__review_reply_start`, or `oracle-pro` review call following `../shared-references/review-tracing.md`. Record the reviewer route, saved threadId, prompt summary, raw response path, decisions, and action items. This preserves the Claude mainline Review Tracing semantics while using Codex-native reviewer calls.
## Key Rules
- **Always ask the Claude reviewer for strict, high-rigor feedback** in every review round.
- Send comprehensive context in Round 1 — the external model cannot read your files
- Be honest about weaknesses — hiding them leads to worse feedback
- Push back on criticisms you disagree with, but accept valid ones
- Focus on ACTIONABLE feedback — "what experiment would fix this?"
- Document the completed `threadId` for potential future resumption
- The review document should be self-contained (readable without the conversation)
## Prompt Templates
### For initial review:
"I'm going to present a complete ML research project for your critical review. Please act as a senior ML reviewer (NeurIPS/ICML level)..."
### For experiment design:
"Please design the minimal additional experiment package that gives the highest acceptance lift per GPU week. Our compute: [describe]. Be very specific about configurations."
### For paper structure:
"Please turn this into a concrete paper outline with section-by-section claims and figure plan."
### For claims matrix:
"Please give me a results-to-claims matrix: what claim is allowed under each possible outcome of experiments X and Y?"
### For mock review:
"Please write a mock NeurIPS review with: Summary, Strengths, Weaknesses, Questions for Authors, Score, Confidence, and What Would Move Toward Accept."