llms.txt · git:20260913.474e464 · 2026-09-13 · sha256 62d7be129e6cca28
llms.txt git:20260913.474e464A
Immutable. This exact content is served forever at /api/v1/blob/62d7be129e6cca28.
# hermes-jailbench Deterministic jailbreak regression benchmark for LLM endpoints. Use this repo when you need to: - rerun known jailbreak attempts after a prompt or model change - compare refusal rates across releases - save markdown or JSON evidence for audit trails Primary CLI: - `hermes-jailbench --demo` - `hermes-jailbench --dry-run` - `hermes-jailbench --model <model> --api-key $ANTHROPIC_API_KEY` Agent skill (one canonical skills/hermes-jailbench/SKILL.md): - Claude Code: `claude plugin marketplace add hermes-labs-ai/hermes-jailbench` then `claude plugin install hermes-jailbench@hermes-jailbench` - Codex CLI: `codex plugin marketplace add hermes-labs-ai/hermes-jailbench` then `codex plugin add hermes-jailbench@hermes-jailbench` - Gemini CLI: `gemini extensions install https://github.com/hermes-labs-ai/hermes-jailbench --ref main` Outputs: - per-attack verdicts - refusal, partial, and compliance counts - category breakdown - markdown or JSON report Do not use this repo as: - proof against novel jailbreaks - a multi-turn red-team harness - a semantic judge for edge cases without human review Key success condition: - the same attack corpus and response text produce the same score and report shape ## About Hermes Labs Hermes Labs is an independent AI-reliability lab building open-source tools that catch silent failure modes in production AI. More at https://hermes-labs.ai