AGENTS.md · diff
git:20260818.60985ae to git:20260910.331ccea
21 added, 0 removed. Audit B to B.
# AGENTS.md
+ <!-- Instruction contract v1.0 — 2026-09-09 -->
+
+ Priority order: preserve security and coverage semantics; preserve public API behavior;
+ then minimize the diff. Treat each user request as an independent task and carry prior
+ task state forward only when the user explicitly asks.
+
Little Canary is a prompt-injection detection library that uses a sacrificial canary model as an inbound risk sensor.
## Use it for
- screening untrusted text before it reaches a main model or agent
- combining structural pattern checks with behavioral compromise checks
- returning `block`, `flag`, `pass`, or `degraded` state plus advisory text
## Do not use it for
- formal security guarantees
- audited benchmark comparisons
- replacing runtime containment or outbound tool controls
+ ## Key paths
+
+ - `little_canary/` — package source
+ - `tests/` — pytest suite
+ - `benchmarks/` — false-positive/red-team runners, not part of the default CLI flow
+ - `.hermes/` — declared quality-gate config run by CI's Hermes quality rail
+
## Minimal commands
```bash
pip install -e ".[dev]"
little-canary --version
little-canary demo --replay # exits 2 until an admitted fixture is packaged
little-canary demo --live --backend ollama --model qwen2.5:1.5b --endpoint http://127.0.0.1:11434
little-canary serve --help
pytest -q
ruff check little_canary tests
mypy little_canary
+ python -m build
```
## Output shape
- Python API separates routing (`safe`) from coverage (`degraded`, `canary_status`, `analysis_method`, `analysis_status`)
- failed/unavailable coverage keeps risk unset and never means an exercised PASS
- CLI exposes explicit replay/live demo paths and the local HTTP server
- benchmark scripts live under `benchmarks/` and are not part of the default CLI flow
## Success means
- when an admitted fixture is packaged, replay reproduces analyzer behavior for those exact bytes without a network call
- a build without that fixture must report `REPLAY UNAVAILABLE`; this is a hold, not a safe result
- live evidence names its backend, model digest, configuration, and observed result
- structural and behavioral layers agree with the documented modes
- the Ollama path and OpenAI-compatible protocol adapter pass their declared
offline tests; live support remains bound to endpoint-specific evidence
## Common failure cases
- the canary backend is unavailable and fail-open routing passes through with a visible degraded state
- users expect this tool to replace broader agent runtime controls
- benchmark claims are quoted without the methodology caveats in the README
## Maintainer notes
- preserve fail-open behavior unless there is an explicit versioned policy change
- never treat replay, mock, or static evidence as a current live-model result
- keep benchmark caveats aligned with README claims
- keep tests offline and mock network calls
+
+ ## Definition of done
+
+ - `pytest -q` and `ruff check little_canary tests` pass; `mypy little_canary` is diagnostic
+ - new behavior has an offline, mocked-network test
+ - fail-open behavior and coverage-state semantics are unchanged unless explicitly intended
+ - README and this file stay consistent with observed CLI/API behavior
## Optional Hermeneutic Gate
For a consequential semantic review of an assistant-generated English draft,
you may run the standalone deterministic second-opinion check from the
separate [Hermeneutic Gate](https://github.com/hermes-labs-ai/hermeneutic#epistemic-gate)
CLI. It is not part of Little Canary's dev extras; install it separately:
```bash
python -m pip install hermeneutic
hermeneutic gate --draft review-summary.md
```
It flags surface shapes such as completion overclaiming, unsupported numeric
claims, relayed authority, and unhedged certainty. It runs offline, does not
invoke Little Canary or a model backend, and does not create a release or
perimeter gate. Little Canary has no repository-level Gate workflow or hook;
callers choose whether and how to enforce the exit code. Per the
[quick start](https://github.com/hermes-labs-ai/hermeneutic#quick-start),
low-severity `RISK` is advisory (exit 0); medium/high `RISK` exits 1. If it
flags a claim, add direct evidence, hedge the claim, or remove the
unverifiable wording.