agentscope-ai/openjudge

21 agent markdown files indexed from agentscope-ai/openjudge. Install any of them pinned to an exact hash with mdr add agentscope-ai/openjudge/<name>.

00-academic-router skill
agentscope-ai/openjudge · skills/academic-eval/00-academic-router/SKILL.md · Use when the user wants help with academic papers or citations but it's unclear which specific workflow fits —…
git:20260911.d1e0642 · audit A · 851 stars
01-paper-review skill
agentscope-ai/openjudge · skills/academic-eval/01-paper-review/SKILL.md · Review academic papers for correctness, quality, and novelty using OpenJudge's multi-stage pipeline. Supports PDF files…
git:20260911.d1e0642 · audit A · 851 stars
02-bib-verify skill
agentscope-ai/openjudge · skills/academic-eval/02-bib-verify/SKILL.md · Verify a BibTeX file for hallucinated or fabricated references by cross-checking every entry against CrossRef, arXiv…
git:20260911.d1e0642 · audit A · 851 stars
03-ref-hallucination-arena skill
agentscope-ai/openjudge · skills/academic-eval/03-ref-hallucination-arena/SKILL.md · Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and…
git:20260911.d1e0642 · audit A · 851 stars
00-arena-router skill
agentscope-ai/openjudge · skills/arena-eval/00-arena-router/SKILL.md · Use when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific…
git:20260911.d1e0642 · audit A · 851 stars
01-auto-arena skill
agentscope-ai/openjudge · skills/arena-eval/01-auto-arena/SKILL.md · Automatically evaluate and compare multiple AI models or agents without pre-existing test data. Generates test queries…
git:20260911.d1e0642 · audit A · 851 stars
02-ref-hallucination-arena skill
agentscope-ai/openjudge · skills/arena-eval/02-ref-hallucination-arena/SKILL.md · Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and…
git:20260911.d1e0642 · audit A · 851 stars
claude-authenticity skill
agentscope-ai/openjudge · skills/claude-authenticity/SKILL.md · Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted…
git:20260310.56772fa · audit A · 851 stars
meta-eval skill
agentscope-ai/openjudge · skills/eval_pipeline/00-meta-eval/SKILL.md · Use when the user wants to build an evaluation system for an LLM/agent application but doesn't know where to start —…
git:20260708.2151def · audit A · 851 stars
eval-design skill
agentscope-ai/openjudge · skills/eval_pipeline/01-eval-design/SKILL.md · Use when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial…
git:20260708.2151def · audit B · 851 stars
metric-design skill
agentscope-ai/openjudge · skills/eval_pipeline/02-metric-design/SKILL.md · Use when the user has evaluation principles or a dataset but needs help choosing the right graders, designing…
git:20260708.2151def · audit A · 851 stars
align-human skill
agentscope-ai/openjudge · skills/eval_pipeline/03-align-human/SKILL.md · Use when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with…
git:20260708.2151def · audit A · 851 stars
eval-report skill
agentscope-ai/openjudge · skills/eval_pipeline/04-eval-report/SKILL.md · Use when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment…
git:20260708.2151def · audit A · 851 stars
rag-eval skill
agentscope-ai/openjudge · skills/eval_pipeline/05-rag-eval/SKILL.md · Use when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating…
git:20260708.2151def · audit A · 851 stars
prompt-regression skill
agentscope-ai/openjudge · skills/eval_pipeline/06-prompt-regression/SKILL.md · Use when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether…
git:20260708.2151def · audit A · 851 stars
redteam skill
agentscope-ai/openjudge · skills/eval_pipeline/07-redteam/SKILL.md · Use when the user wants to test their LLM/agent application for safety and security vulnerabilities — jailbreaks…
git:20260708.2151def · audit A · 851 stars
bootstrap skill
agentscope-ai/openjudge · skills/eval_pipeline/08-bootstrap/SKILL.md · Use when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch…
git:20260708.2151def · audit A · 851 stars
find-skills-combo skill
agentscope-ai/openjudge · skills/find-skills-combo/SKILL.md · Discover and recommend **combinations** of agent skills to complete complex, multi-faceted tasks. Provides two…
git:20260309.343828b · audit F · 851 stars
mmx-cli skill
agentscope-ai/openjudge · skills/mmx-cli/SKILL.md · Generate text, images, video, speech, and music via the MiniMax AI platform. Covers text generation (MiniMax-M3 model)…
git:20260609.6420ce5 · audit A · 851 stars
01-graders-and-pipeline skill
agentscope-ai/openjudge · skills/openjudge-core/01-graders-and-pipeline/SKILL.md · Build custom LLM evaluation pipelines using the OpenJudge framework. Covers selecting and configuring graders…
git:20260911.d1e0642 · audit A · 851 stars
02-rl-reward skill
agentscope-ai/openjudge · skills/openjudge-core/02-rl-reward/SKILL.md · Build RL reward signals using the OpenJudge framework. Covers choosing between pointwise and pairwise reward strategies…
git:20260911.d1e0642 · audit A · 851 stars

Source

github.com/agentscope-ai/openjudge · 851 stars · license Apache-2.0 · pushed 2026-09-11