Immutable. This exact content is served forever at /api/v1/blob/8ffe554ff3355378.
---
name: agent-smith
description: >-
Offload bulky or research-heavy text work to the already-paid fleet — z.ai GLM for bulk
drafting, Google Gemini (via GEMINI_API_KEY) for --file and --search — then
verify and finish with Claude to spare Claude's tokens and context. Trigger PROACTIVELY,
even if Gemini isn't named, whenever the heavy part of a task is generating or digesting
text: web research on current facts or "what's new/changed in X" with source links;
summarizing, digesting, or extracting from one long document OR many files (PDFs,
transcripts, interviews, logs, CSVs, contracts) into bullets, tables, quotes, or themes;
classifying or transforming many records; drafting a plan, proposal, roadmap, or research
write-up; website, landing-page, or marketing copy (blog, FAQ, headlines, first-draft
HTML/CSS); config or infrastructure boilerplate as text (wrangler.toml, Worker scaffolds,
Dockerfiles, CI YAML, IaC modules); turning one announcement into many platform-specific
social or marketing posts; or any mass first-draft code or boilerplate. Trigger for CODE
work too: drafting a new module, class, CLI tool, or test suite from a clear spec; porting
or translating code; or a multi-step scratch-sandbox build (fix-a-bug, add-a-feature,
build-a-small-app) via the bundled smith_agent.py tool loop. Also trigger when the user
says "use Gemini" or "agent-smith", mentions the local fleet, asks to save Claude tokens,
or asks where tokens go / to audit or reduce Claude token usage. Gemini drafts; Claude scopes,
cross-checks, integrates, and delivers. Do NOT use for short/quick/interactive work, small
edits, correctness-critical debugging, or security-sensitive tasks — and never for the
EXECUTION half of a task: deploying to Cloudflare, posting to a live account, committing or
building, or anything needing your credentials, tools, or live repo context. Gemini has no
access to your accounts or tools, so those stay on Claude or a script you control. Don't use
it to install or configure Gemini itself.
---
# Gemini Offload
## The deal
You (Claude) are the orchestrator and the quality bar. **Delegate the bulk, keep the
judgment.** Backend output is an intermediate, never a deliverable — pull every result
through your own review before it counts. This spends the fleet's tokens instead of
Claude's, per the user's standing quota-economy preference.
## What to offload vs. keep
**Offload** (voluminous and checkable): research/fact-finding (`--search`), digesting big
inputs (`--file` — let the backend read the 200 pages, not you), first-draft generation
(code, configs, tests, copy from a clear spec).
**Keep** (subtle, stateful, expensive to get wrong): orchestration, repo-aware edits,
correctness-critical reasoning, security, final review + integration.
**Break-even:** don't offload small work — overhead exceeds savings below ~50 KB input /
a page or two of output (measured; see [references/measured-results.md](references/measured-results.md)).
**Route elsewhere — persistent corpus:** when the corpus is one Josh will interrogate
*repeatedly* (not one-shot digested), or he wants an audio overview of it, use the
`notebooklm` skill instead of `--file`/`long-digest` — a notebook persists and re-answers,
a `gemini.py` call is one-and-done. ⚠ NotebookLM runs do NOT touch `data/usage.jsonl`, so
`gap_report.py` cannot see them — say so in the report or the route reads as "Claude did it".
## Task playbooks
**Offload the words, keep the action** — anything that deploys, commits, posts, or charges
stays with you (backends have no credentials).
| Workflow | Offload | Keep |
|---|---|---|
| Planning | research, draft plans, option write-ups | the decision, repo specifics |
| Website content | copy, blog/FAQ, first-draft HTML/CSS | wiring, voice/legal pass, browser test |
| Cloudflare/infra | wrangler.toml, Dockerfiles, CI YAML *as text* | the deploy — your MCP + creds |
| Business postings | announcement → per-platform posts | actual posting; final approval |
| Info-doc HTML (the HTML-docs rule) | md-draft → styled HTML via `references/html-doc-shell.html` | the draft's content, parity check, saving |
Full recipes: [references/playbooks.md](references/playbooks.md).
## The loop
1. **Scope** — tight, self-contained prompt; the backend has none of this conversation.
2. **Delegate** — pick backend + model below.
3. **Review in proportion to payload** — backends hallucinate APIs/citations; run/lint code.
**Never re-ingest a large input to verify** — sample: check format, a few known anchors,
spot-check sections. (Full re-read = you paid twice.)
4. **Finish** — integrate yourself; send focused revision prompts rather than redoing.
5. **Report + record the verdict** — tell the user what you delegated and verified, then:
`python3 "$SKILL/scripts/verdict.py" good` (or `bad "why"`). One line, every review.
## Using the helper
`scripts/gemini.py` — answer on stdout; model/tokens/sources on stderr (read both).
Pure stdlib, runs on macOS/Linux/Windows (`python3` vs `python`; skill dir
`~/.claude/skills/agent-smith` vs `%USERPROFILE%\.claude\skills\agent-smith`).
Key: `GEMINI_API_KEY` (fallback `GOOGLE_API_KEY`) in the environment.
**`--tag SHAPE` is REQUIRED on every call** (enforced in code 2026-07-28 — an untagged call
exits 2 before spending anything). Habit alone left 863/915 runs untagged and therefore
invisible to `gap_report.py` and the hebbian router. Reuse an existing shape or coin one;
use `--tag smoke` for throwaway routing checks so they stay out of the review queue.
```bash
SKILL=~/.claude/skills/agent-smith
python3 "$SKILL/scripts/gemini.py" --tag copy-draft "Explain X in 5 bullets"
cat spec.txt | python3 "$SKILL/scripts/gemini.py" --tag draft-spec "Make a checklist"
python3 "$SKILL/scripts/gemini.py" --search --tag research "What's new in Swift?"
python3 "$SKILL/scripts/gemini.py" --file report.pdf --tag long-digest "Summarize as bullets"
python3 "$SKILL/scripts/gemini.py" --file inv.pdf --schema s.json --tag classify "Extract items"
# BATCH (local, zero Claude tokens per item): manifest = one path per line;
# images ride vision, text appends to prompt; per-item .out.txt + ONE JSON summary.
python3 "$SKILL/scripts/gemini.py" --backend ollama --batch files.txt --tag classify --out-dir out "Classify: ..."
# CONSENSUS batch (disagreement fires escalation): each item on TWO local models at
# temp 0; agree -> accept; disagree -> .A/.B files + _escalate.txt queue. SHORT outputs only.
python3 "$SKILL/scripts/gemini.py" --backend ollama --model llama3.2:3b \
--batch records.txt --consensus gpt-oss:20b --tag classify --out-dir cls "One word: ..."
```
Windows: same flags, `python` launcher, `Get-Content` for stdin.
**Flags:** `--tag` (REQUIRED) · `--model` · `--system` · `--file` (repeatable) · `--search` ·
`--json`/`--schema` · `--temperature` · `--max-tokens` · `--thinking-budget N` ·
`--preflight` (gemini-cli syntax check) · `--list-models` · `--no-tailor` (below).
API internals: [references/gemini-api.md](references/gemini-api.md).
**Per-model prompt tailoring (2026-07-26).** `--system` carries TASK framing (role/style);
`gemini.py` separately auto-appends a MODEL framing clause — a known, measured failure mode
for whichever model is actually resolved (e.g. gpt-oss:20b → strip reasoning-residue dead
code before returning; qwen3-vl:4b → flag long digit strings as unverified unless reading a
crop). Every call through a profiled model gets the fix instead of relying on Claude to
retype a lane note by hand. Table = `MODEL_PROFILES` in `gemini.py` (evidence-only: add an
entry after a `verdict.py bad`, not a hunch). `--no-tailor` opts a single call out; a
`[tailor]` stderr line + `"tailored"` ledger field confirm it fired. `ROUTE_WARNINGS` (not
suppressed by `--no-tailor`) separately flags a known-bad model+tag combo already called out
as a Route BLOCK below — a routing decision, not something a prompt clause should paper
over. Design + verification: [references/model-tailoring-2026-07-26.md](references/model-tailoring-2026-07-26.md).
## Backends (`--backend`) — default `gemini`, reach for edges
| backend | runs on | cost | files/web | use when |
|---|---|---|---|---|
| `gemini` | Google cloud (`GEMINI_API_KEY`) | free tier, rate-limited | **yes** | ⚠ **specialist since 2026-09-05, no longer the bulk default** — reach for it for `--file`/`--search`, the only things ONLY it can do. Its quota is the **API's**, separate from the consumer Gemini Pro plan; Pro buys this backend nothing |
| `gemini-cli` | your OAuth login | **PAID — Josh's Google AI Pro, already bought** | no | THE lane when the free API is slow/429ing (measured 2026-07-12: 3.8s vs 50s+ congested API). This is the ONLY backend that spends the Pro plan — `--backend gemini` does not. Text-only: pipe file contents via stdin. Use what's paid for |
| `fm` | this Mac (~3B) | free | no | private + simple bulk (`FM_HELPER` path) |
| `ollama` | this Mac | free, unlimited | images yes | private/offline/high-volume; the FLEET below |
| `openai` | any OpenAI-compatible URL | free tiers exist | no | burst beyond Gemini; shorthands `groq`\|`openrouter`\|`openai`\|`ollama`; auth `OPENAI_API_KEY` (Groq: `GROQ_API_KEY`). Groq `openai/gpt-oss-120b` = verified free frontier-adjacent. **Free clouds may train on your data — private work stays local** |
## Local fleet routing (gym-earned; evidence → [references/measured-results.md](references/measured-results.md))
- **Quality code / app builds:** `gpt-oss:20b` (12 GB) — TRUSTED code-gen/struct/edits/
app-builds (double perfect sweep). Review watch-item: reasoning residue (commented-out
debug prints, dead branches, doc claims for absent code).
- **Vision + design: `gemma4:26b` REMOVED 2026-08-16** (18 GB reclaimed, Josh's call). It held
the design crown and was the image auto-pick, but had gone 4 weeks unused and the vision-v1
suite measured it INVENTING 4 fields on a tall page (20% recall in 6s — it wasn't reading)
where `qwen3-vl:4b` scored 100% at a fifth the size. **`gemini.py`'s image auto-pick now
resolves to `qwen3-vl:4b`.**
- **DESIGN lane RE-GATED 2026-08-16 — `gpt-oss:20b` takes it (20.0/24 blinded).** 4-way
re-gate after gemma4 left: gemini-pro 22.5 · **gpt-oss:20b 20.0** · qwen3-coder:30b 15.0 ·
gemini-flash incomplete (503). So local design work goes to `gpt-oss:20b`, **not**
`qwen3-coder:30b`, which is weak here (dead branches, falsely-documented exceptions, a
printed success message for work it never did) — it holds the SPEED lane, not this one.
⚠ Known, REPLICATED defect: gpt-oss:20b writes *int-only `consume` on a float bucket* in
rate-limiter-shaped code — same flaw as 07-04, six weeks apart. Review numeric-boundary
code from it. ⚠ `--tag design` still defaults to **flash**, deliberately: flash TIED pro
(11.5) on the only task both finished, so there is no measured basis for a pro override.
Pass `--model pro` by hand when a design task deserves the ceiling.
- **Vision pre-screen — TRUSTED (2026-07-12, 2-consecutive gate):** `qwen3-vl:4b` (3.3 GB
dl, ~8 GB loaded, 256k ctx) — co-resides with gpt-oss:20b. Run 1: **9/9** (incl. tiny-text
OCR on a dense SC editor); run 2: **8/9 on fresh corpus incl. a TALL 5265px scroll** where
it read even small-text prices correctly (no gemma4-style invention). The one miss —
dropped a leading digit in a 10-digit app ID on a full-window shot — read EXACTLY on a
field crop. **Lane rule: for exact long digit strings (IDs, serials, keys), crop the field
first or double-read.** Use for "which screen / did the dialog open / read this field":
`--backend ollama --model qwen3-vl:4b --file shot.png`. gemma4:26b keeps the
quality/design crown.
- **Fast bulk drafts:** `qwen3-coder:30b` (18 GB) — ollama default; 2–8s one-shots.
- **Long private digests — NEW LANE (validated 2026-07-12):** `gpt-oss:20b` at up to
**131k context** — RAM stays flat at 12 GB (MXFP4 MoE), 3/3 needle recall + correct
comprehension measured at 52k tokens (~4.6 min). `gemini.py` now auto-sizes `num_ctx`
from input length (Ollama silently truncates otherwise). Use for transcripts/contracts/
logs too private for free cloud tiers; beyond ~130k tokens split it or use `--backend
gemini` (1M). Caveat: the witness re-run doubles the cost of a long-prompt call.
- **Lighter backup: REMOVED 2026-08-16.** `qwen2.5-coder-smith:14b` (9 GB, our gym-gated
fine-tune) was deleted after 18 ledger runs and zero good/bad verdicts in either
direction — an unverified route occupying disk, not a trusted one. Re-pull and re-earn
a slot with real verdicts before routing to it again.
- **Bench / second opinion:** `agents-a1` (21 GB, trusted everywhere, no lane) — decorrelated
lineage; premium consensus/witness third voice. `llama3.2:3b` = tiny text floor only.
- **Cloud model choice:** `flash` for bulk text; **`pro` for code/design/research synthesis**
and as escalation when local attempts fail.
- **Residency (36 GB Mac, since the claude-mem observer went local 2026-07-12):** the observer
keeps `gpt-oss:20b` (12 GB) hot most of the day, and it + `qwen3-coder:30b` (18 GB) can't
co-reside in GPU memory — routine one-shots on the 26b/30b now pay a 20–60s swap and evict
the observer's model. Prefer `--model gpt-oss:20b` for routine local drafts; reach for
gemma4:26b (vision) / qwen3-coder:30b deliberately and expect the swap. If Ollama wedges
(model stuck "Stopping...", requests hang): `kill` the `llama-server` runner PID, or
restart Ollama.app.
- Always: **the model drafts, you verify** — every winner has shipped a bug a review caught.
## DEFAULT-TO-LOCAL / DEFAULT-TO-PAID routes — gap report 2026-07-25, updated 2026-07-28
**Enforced in code as of 2026-07-27 (cost pass — user wants minimum spend):** `gemini.py`'s
`DEFAULT_LOCAL_FOR_TAG` table auto-routes `--tag doc-format|classify|vision-prescreen|
subagent-fanout` to the free local model below whenever BOTH `--backend` and `--model` are
left unset — logs `[cost] --tag X defaults to local`. Passing either flag explicitly still
wins (e.g. `--backend gemini` forces cloud). `--search` always forces cloud (no local web
grounding) regardless of tag.
**`code-draft` AND `long-digest` moved to a NEW `DEFAULT_PAID_FOR_TAG` table as of
2026-07-28** — bare `--tag code-draft` or `--tag long-digest` now route to the z.ai
**GLM Coding Plan** (`glm-5.3`, `https://api.z.ai/api/coding/paas/v4`, needs
`ZAI_API_KEY`), logging `[paid] --tag X defaults to the z.ai GLM Coding Plan`. Flat $18/mo
subscription — the marginal cost of one more call is $0, so this table is checked BEFORE
the free-local table and wins where both would apply. `code-draft`'s justification:
agent-gym swept it 27/27 the same day (all 5 agentic tasks incl. native tool-calling), beat
the local TRUSTED baseline's wall-clock on 4/5 agentic tasks, and the user explicitly asked
to prioritize a resource already paid for.
`long-digest`'s justification is different and worth stating precisely, because it's a
judgment call, not just a benchmark win: a head-to-head vs the local TRUSTED baseline
(`gpt-oss:20b`) on the same document + same 5 questions scored GLM 4/5 in 25s vs
gpt-oss:20b 5/5 in **4:49** — an 11x wall-clock gap for a one-point accuracy difference.
The user's call: Claude verifies every delegated output regardless of which model drafted
it — that's not new for this tag, it's the standing rule — so a small accuracy gap
shouldn't outweigh an 11x time cost; verification is the constant, optimize the other
variable. GLM's context window (1M, docs-verified) comfortably exceeds gpt-oss:20b's 131k,
so this isn't a capacity tradeoff either. **What this does NOT resolve: privacy.**
Verification catches wrong facts in a cloud draft; it does not catch data having already
left the machine. For anything genuinely sensitive (credentials, PII, private content),
route explicitly to local — `--backend ollama --model gpt-oss:20b` — by hand, every time;
the paid default does not know to make that call for you.
Every paid-default tag auto-raises `--max-tokens` to a measured floor (**32000 since
2026-09-05**, was 8000) if the caller left it unset or too low — a budget calibrated for one
shape can leave another shape returning EMPTY content, since GLM spends part of its budget on
a hidden reasoning channel before writing anything visible. Passing `--max-tokens` explicitly
still wins if it's already at or above the floor.
**Why 32000, measured against this exact endpoint 2026-09-05:** glm-5.3 returned
`finish_reason=length` with `content=""` on **six of six** attempts at `max_tokens` 8000,
spending 7994–8000 of its 8000 tokens on the hidden reasoning channel, then wrote a full
19–25 KB module on **three of three** at 32000 (13.6k–23.2k reasoning tokens first). 8000
was never a floor, it was a coin flip that had been landing. Raising a **ceiling costs
nothing** — you are billed for what is generated, not for what you allowed — so treat this
as insurance, not spend. The table's model string also moved `glm-5.2` → `glm-5.3`: z.ai
retired 5.2 in place on 2026-08-14, so the old string had been silently running 5.3 while
logging the wrong name into the ledger.
`--backend ollama --model qwen3-coder:30b` (or `gemma4:26b`) still works as an explicit
override for `code-draft` — local stays free/private, just no longer the silent default.
Any remaining tag with neither a `DEFAULT_PAID_FOR_TAG` nor `DEFAULT_LOCAL_FOR_TAG` entry
still defaults to the Gemini API — their evidence isn't strong enough yet to force a
different route silently.
Reach for these WITHOUT re-deriving the gap report; each is a measured gap joined with a
ledger-trusted (or trial-ready) route. Tag every run so the streak builds.
- **Read-only fan-out / grep-and-summarize sweeps** → `--backend ollama --batch`
(`--tag subagent-fanout` — must match `gap_report.py`'s `SHAPES` tag exactly, or the run
is invisible to the gap report; `fanout-digest` was a stale/wrong tag name here until
2026-07-27, never actually used). Biggest untouched vein: 663 Claude-side Agent/Explore
calls, still **0% delegated** as of 2026-07-28. First real evidence landed that day
though: a 10-file read+summarize test against `zai-glm-5.2` scored 10/10 against ground
truth (verdict good) — a genuine trial-ready result, not just an untested route anymore.
**Correctness-critical work (code review, bug hunts) is fair game for this tag too** —
the fleet drafts candidate findings, Claude re-verifies every one against the real file
before acting; see playbook 6 in
[references/playbooks.md](references/playbooks.md) for the prompt template that fixed a
measured 0/4-false-positive run (`--max-tokens` floor raised to 8192, `gpt-oss:20b`'s
profile clause now demands a quoted line + restated guard before any finding).
- **Long-document digest** (logs, transcripts, CSVs, contracts) → bare `--tag long-digest`
now defaults to the paid GLM Coding Plan (see DEFAULT-TO-PAID above). Route explicitly to
`--backend ollama --model gpt-oss:20b` (131k ctx) for anything sensitive — that's a
privacy call the default can't make for you.
- **Screenshot / vision prescreen** → `--backend ollama --model qwen3-vl:4b --file shot.png`
(`--tag vision-prescreen`; 3g/0b TRUSTED). ~127M Claude context tokens of screenshots;
local first-pass describes, Claude views only flagged shots. Digit-string crop rule applies.
- **Mechanical code boilerplate from a clear spec** (Codable conformance, enum plumbing,
UI-label tables, test scaffolds) → bare `--tag code-draft` now defaults to the z.ai GLM
Coding Plan (`glm-5.3`, see the DEFAULT-TO-PAID section above) — gym-swept 27/27
2026-07-28, no residency swap. `--backend ollama --model qwen3-coder:30b` (or
`gemma4:26b`) still available as an explicit local/private override.
- **Classification / tag / label batches** → `llama3.2:3b` (`--tag classify`; 1g/0b),
`--consensus gpt-oss:20b` when accuracy matters.
- **Web research** → `--search --tag research`. `gemini.py` now defaults `--tag research`
to **pro** automatically when `--model` is omitted (fixed 2026-07-27 — was silently
defaulting to flash, which scored 0g/1b, MISROUTED; pro scores 9g/1b). Passing
`--model flash` still works but logs a route warning. The remaining gap is volume, not
model choice: 2% delegated as of 07-25 — the blocker is habit, not tooling.
- **Research on z.ai — SECOND LANE, NOT A REPLACEMENT (measured 2026-08-10).** z.ai's own
`web_search` tool works on the Coding Plan endpoint and is wired to `--search` on the
`openai` backend (z.ai hosts only; `SEARCH_MIN_TOKENS=8000` floor, since grounded runs
put ~89% of the completion budget into hidden `reasoning_content`). But on a 10-question
current-release eval with ground truth read from primary sources
(`evals/research_grounding_eval.py`): **gemini-pro 10/10, glm-5.2 + web_search 7/10.**
Every GLM miss was the same failure — *stale pre-training fact, real citations attached*
("Python 3.14.0, released October 7 2025 [2,3]" vs the true 3.14.7 / 2026-08-05).
Retrieval was fine (5.0 sources/question vs Gemini's 2.6); it answered around them.
So: **`--tag research` stays on gemini-pro.** z.ai only takes research when Gemini's
monthly cap blows — `gemini.py` now reroutes automatically on a cap-shaped 429 and
prints an UNVERIFIED banner; re-verify every version and date it returns.
`SMITH_NO_FALLBACK=1` disables the reroute. Transient 429s are NOT treated as caps.
- **Route BLOCK:** `doc-format @ gpt-oss:20b` is 0g/**3b** — do not send doc-format there
until it passes an agent-gym task. Use gemini-pro (5-streak, LIGHT REVIEW — first earned
tier), or gemma4:26b / qwen3-coder:30b (2g/0b each).
## Gemini is no longer the default for bulk — 2026-09-05
**`translate` and `copy-draft` joined `DEFAULT_PAID_FOR_TAG`**, so bulk drafting now routes
to z.ai by default rather than to Gemini. Two facts drive it, and the first is the one that
is easy to get wrong:
⚠ **The consumer Gemini Pro subscription does NOT raise this script's quota.** They are
different products. `gemini.google.com` (where Josh is on **Pro**) is the consumer app;
`--backend gemini` calls `generativelanguage.googleapis.com` with `GEMINI_API_KEY`, which is
the **API**, on its own separate quota. Burning the consumer allowance has no effect here,
and Pro buys this script nothing. Josh declined the extra ~$25/mo paid API tier, so the API
stays on its **free, rate-limited** tier — which is exactly what makes it the wrong default
for volume.
The second: z.ai is already paid for, its marginal cost is $0, and on the only head-to-head
measured it is also the better option (TRANSLATE 83% vs Ollama Cloud's 60%; it emitted a full
module at 32000 where ollama.com's `/v1` returned nothing).
**Gemini keeps what only it can do:** `--file` and `--search`. Nothing else routes there by
default. Deliberately NOT moved: `design` stays on **flash** (flash TIED pro 11.5 on the only
task both finished — no measured basis to move it) and `research` stays on Gemini because
`--search` forces cloud and no other backend has web grounding.
⚠ **z.ai limits CONCURRENCY, not volume:** ~6 requests in flight, then HTTP 429
`code 1302, Rate limit reached for requests` fires instantly. Keep fan-outs serial-ish.
**`--backend gemini-cli` is a separate, already-paid lane** — it runs on Josh's OAuth login
against the Pro subscription above, so it spends something already bought rather than the
free API tier. Reach for it when the free API is slow or 429ing. Text-only: pipe file
contents via stdin.
## Standing offload targets — token audit 2026-07-12
Measured: 136M Claude output tokens/30d, only 0.2% offloaded. Biggest single consumer
(claude-mem observer, ~16% of ALL output) now runs on local `gpt-oss:20b` via claude-mem's
openrouter provider → `http://127.0.0.1:11434/v1`. Ranked remaining targets — main-loop web
research (habit gap: 679 searches + 701 fetches ran on Opus), read-only subagent fan-outs,
first-draft code via smith_agent, screenshot pre-screening, session hygiene — with numbers,
the claude-mem config/revert path, and the generalizable "point any custom-endpoint tool at
the fleet" precedent: [references/token-audit-2026-07-12.md](references/token-audit-2026-07-12.md).
Re-run the audit: `python3 /Users/joshualangberg/Python/docs/token_audit.py`.
## Agentic offload — smith_agent.py (sandboxed tool loop)
Multi-step repo tasks (fix a bug, add a feature, build a small app). `--backend`/`--model`
are explicit flags every time (smith_agent.py isn't tag-routed like gemini.py), so this
is a documented DEFAULT PRACTICE, not something the tool enforces — reach for it by hand:
```bash
# DEFAULT since 2026-07-28: z.ai GLM Coding Plan — gym-swept 27/27 incl. APP-BUILD 3/3
# and CODE-EDIT 2/2 at L3 (real tool loop), beat local gpt-oss:20b's wall-clock on 4/5
# agentic tasks, flat-rate subscription (marginal cost $0). Needs ZAI_API_KEY.
python3 "$SKILL/scripts/smith_agent.py" --backend openai \
--base-url https://api.z.ai/api/coding/paas/v4 --model glm-5.3 \
--api-key-env ZAI_API_KEY --workdir /path/to/SCRATCH --prompt-file task.txt
# local/private fallback (free, no data leaves the machine):
python3 "$SKILL/scripts/smith_agent.py" --model gpt-oss:20b \
--workdir /path/to/SCRATCH --prompt-file task.txt
# big single-file writes: add --max-gen-tokens 4096 (default 1600 truncates them)
# cloud escalation: --backend gemini --model pro (5x slower, 503 risk)
```
Rules: SCRATCH dirs only (it executes model shell — never a live repo); write the task like
a ticket (spec, exact outputs, how to verify); seed a `test_public.py`; **verify the result
yourself**, then verdict it. Canonical source + harness: `~/Python/agent-gym/`.
## Local transcription — transcribe.py (audio → text, free, private)
`python3 "$SKILL/scripts/transcribe.py" FILE` (wav/mp3/m4a/aiff; `--timestamps`;
`--model small` for speed). mlx-whisper on Apple Silicon, ~2–4s per clip after load,
~98% semantic accuracy (measured). **Caveat: verify rare proper nouns by eye** (Sanskrit/
domain terms can be misheard). Pattern: transcribe locally, then offload the text digest.
## Progress tracking — ledger, verdicts, routing weights, witness
- Every run appends one JSON line to `data/usage.jsonl` (fail-safe; `SMITH_LEDGER` overrides).
Review: `python3 "$SKILL/scripts/usage_report.py"` (`--today`, `--last N`, `--unreviewed`).
- **Say what a run is FOR (2026-07-18).** The ledger used to record only size/speed/model, so
887 runs were indistinguishable after the fact and nothing could be reviewed or routed on
purpose. Records now carry `purpose` (auto-derived from the prompt), `tag`, `project`, and
input filenames. **Pass `--tag SHAPE` on every delegation** (`research`, `doc-format`,
`classify`, `vision`, `copy-draft`, `long-digest`…) — it groups the report and feeds the
hebbian routing weights below; add `--purpose "…"` when the first prompt line makes a poor
title. `SMITH_LOG_PROMPTS=0` records only explicit purposes. `usage_report.py --unreviewed`
is the review queue — coverage was ~4% (799/887 unverdicted), so the routing weights rest on
a thin sample; verdict a few whenever you're already in the ledger.
- **Verdicts:** `ok` = completed, not correct. After review:
`verdict.py good|bad|stale "why" [--tag TASKTYPE] [--model M] [--script S]`. Tags feed the
report's **hebbian routing weights**: per (task-shape, model) quality + streak → review
tier (≥5 light review, ≥10 spot-check). Good strengthens a route; one bad resets it.
`stale` = the output no longer exists to judge (pre-archive run, deleted workdir); it clears
the queue WITHOUT counting as good or bad.
- **Outputs are archived (2026-07-28).** Every run writes its answer to
`data/outputs/<date>/<ts>-<script>-<model>.txt` and the ledger row carries `output_file`.
Before this, the loop demanded a verdict on every run while the tool stored nothing to
judge — 17 runs aged into permanently ungradeable. `SMITH_NO_ARCHIVE=1` opts out.
- **The WebSearch nudge (2026-07-28).** `scripts/research_nudge.py` is a non-blocking
`PreToolUse` hook on `WebSearch` (wired in `~/.claude/settings.json`). It stays quiet for
the first 2 searches of a session, then prints the exact `--search --tag research` command
on the 3rd and every 4th after. Tunable via `SMITH_NUDGE_AFTER` / `SMITH_NUDGE_EVERY`.
Verified live: `hookSpecificOutput.additionalContext` DOES reach Claude on `PreToolUse`
(the docs are ambiguous; this was tested, not assumed) and the search still runs.
- **Feed failures back:** every `bad` = a ready-made regression test → new agent-gym task
before that shape is delegated again.
- **Find the gaps — `python3 "$SKILL/scripts/gap_report.py"`.** The ledger only sees work that
WAS delegated; it is structurally blind to work Claude did itself, which is where the gaps
are. This joins the ledger against `token_audit.json` (mined from Claude transcripts) and
splits the result two ways: **UNUSED** (Claude did it while a trusted route sat idle) and
**MISROUTED** (delegated to a model the ledger scores badly at that shape). Standing
measurement (unchanged 07-19 → 07-25): **web research 2% delegated** — 1,467 Claude calls
vs 31 fleet runs while `research @ gemini-pro` sits at 9g/1b.
**⚠️ Its Agent-call count OVERSTATES the fan-out gap — corrected 2026-08-10.** It counts
every `Agent` spawn, but of 482 in 30d, 130 were `code-reviewer` and ~225 were
`subagent-driven-development` writing code ("Implement Task 2: …"). The genuinely
read-only, delegable population is **~100–120, not 482** — so read that row as an upper
bound and judge the raw fleet run count, not the ratio. (Same class of error as the
ratio drift: fleet counts are ledger-lifetime, Claude counts a sliding 30d, so the
percentage climbs on its own as a busy month ages out.)
Ranked routes now live in "DEFAULT-TO-LOCAL routes" above. Run it monthly, on gym day, or whenever quota gets tight;
`--refresh` re-mines the transcripts first. Caveat it prints itself: tags only exist on runs
since 2026-07-18, so pre-tag history is recovered by per-shape heuristics and ratios read as
a floor.
- **Fleet identity check:** `ollama pull` updates weights IN PLACE — a trusted model can
silently become a different model. `python3 "$SKILL/scripts/fleet_check.py"` compares
current digests vs the accepted baseline (`--accept` after a deliberate update + re-gate);
run it when anything smells off, and always after pulling updates.
- **Witness comparison is STRUCTURAL, not textual.** Python uses the stdlib `ast`; other
languages use tree-sitter when installed — `pip install tree-sitter tree-sitter-language-pack`
(OPTIONAL: absent, non-Python falls back to exact-match, the pre-2026-07-20 behaviour, and
nothing raises). Without it the sensor could not tell benign formatting from real drift in
Swift/TS/YAML — spacing, semicolons and comments all read as disagreement, which is what
pinned it at 18% agreement. Operators and keywords are kept in the fingerprint (they are
anonymous tree-sitter nodes and were briefly dropped, making `a + b` equal `a - b`);
delimiters are excluded.
- **Self-healing tool calls:** if Ollama's server-side tool parser 500s mid-session
(truncated/complex calls — the known killer), smith_agent now drops native tool schemas
and continues the SAME session via the JSON-fallback protocol instead of dying.
- **Ollama-down failover:** the whole local fleet also runs through any OpenAI-compatible
server via `--backend openai --base-url` (llama-server, `mlx_lm server` — proven, LM
Studio). Ollama is the hub, not a dependency. LM Studio pilot (install, CLI, perf vs
Ollama, a real tool-loop bug found + fixed): [references/lmstudio-pilot-2026-07-15.md](references/lmstudio-pilot-2026-07-15.md).
- **Remote OpenAI-compatible lanes (`--base-url` shorthands):** `groq` · `openrouter` ·
`openai` · `ollama` · **`zai`** · **`cloudflare`**. Auth resolves per host —
`GROQ_API_KEY` · `ZAI_API_KEY` · `CF_API_TOKEN` · else `OPENAI_API_KEY`. `cloudflare`
also needs `CF_ACCOUNT_ID` (32-hex; a numeric id is NOT an account id and 404s with
error 7003).
```bash
export ZAI_API_KEY=... # keys live in the env, NEVER in a file
python3 "$SKILL/scripts/gemini.py" --backend openai --base-url zai \
--model glm-4.5-flash --max-tokens 3000 --tag code-draft "..."
```
- **Vision stays LOCAL — bake-off settled 2026-07-28, do not re-litigate.** Screenshot/UI
triage was the biggest measured offload gap (507M chars on Claude), so the Cloudflare
vision models were tested against the incumbent on one full-res app-screenshot crop with
exact ground truth (`Blue Jay 52% / Northern Cardinal 40% / American Robin 8%` — chosen
because digit strings are `qwen3-vl:4b`'s known weak spot):
| model | score | note |
|---|---|---|
| **`qwen3-vl:4b` (local, 3.3 GB)** | **6/6** | every species and digit exact, identical across trials. Free + private |
| `@cf/llava-hf/llava-1.5-7b-hf` | **0/9** | does not OCR — it DESCRIBES AN IMAGINED SCENE ("two birds: a Robin and a Cardinal"), missed Blue Jay entirely, invented 60/40 then 50/50, and hallucinated a "European Robin" when downscaled |
| `@cf/moondream/moondream3.1-9B-A2B` | untested | rejects every documented image shape (`Type mismatch of '/image'`) with no schema published; priced $0.30/$1.00 |
llava's wiring is fine (`{"image":[bytes],"prompt","max_tokens"}` → HTTP 200) — it is a
**capability** failure, not a config one. **A tiny local model beat both remote options
outright.** Vision offload is closed until a genuinely stronger vision model appears.
- **Embeddings + reranking — `scripts/embed.py` (NEW 2026-07-28).** The fleet had neither
until now; every retrieval question was Claude reading files. Free on Workers AI, needs
`CF_API_TOKEN` + `CF_ACCOUNT_ID`.
```bash
# rank a corpus against a query — the high-value one
python3 "$SKILL/scripts/embed.py" --rerank --query "reverse a linked list" \
--docs corpus.txt --top-k 5 --tag rerank
# raw 1024-dim vectors for your own index
python3 "$SKILL/scripts/embed.py" --docs corpus.txt --out vectors.jsonl --tag embed
```
`@cf/baai/bge-m3` (1024-dim, multilingual) · `@cf/baai/bge-reranker-base`.
**Pick by score shape, measured:** the reranker gives a CLIFF (0.99 / 0.97 for two
relevant docs, then 0.02 for the first irrelevant one) so it is the precision filter;
embeddings give a GRADIENT (0.72 relevant, 0.51 loosely related, 0.35 unrelated) so they
are for recall and indexing. Reranking a shortlist beats thresholding cosine similarity.
Chunks at 100 candidates and merges on score — the ids the API returns are chunk-local
and are mapped back to caller indices, so never trust a raw `id` as a corpus index.
**Cloudflare Workers AI — NOTHING is free; there is a small daily allowance
(corrected 2026-07-28, an earlier note in this file wrongly said "FREE").** Every model
is priced. Workers AI bills **$0.011 per 1,000 neurons** with a **10,000 neuron/day free
allocation ≈ $0.11/day**. On **Workers Free the allowance is a HARD STOP** — calls fail
once it is gone; only Workers Paid bills overage. `@cf/zai-org/glm-5.2` is additionally
gated: HTTP 403 code 5035 "requires a Workers Paid plan" regardless of allowance.
What $0.11/day actually buys (derived from the published per-model rates):
| model | in tok/day | out tok/day | verdict |
|---|---|---|---|
| `@cf/baai/bge-reranker-base` | ~35 M | — | effectively unlimited at our scale |
| `@cf/baai/bge-m3` | ~9 M | — | effectively unlimited |
| `@cf/openai/gpt-oss-120b` | ~314 k | ~147 k | real but EXHAUSTIBLE — a few big jobs/day |
| `@cf/qwen/qwen2.5-coder-32b-instruct` | ~167 k | ~110 k | exhaustible |
| `@cf/zai-org/glm-5.2` | ~79 k | ~25 k | blocked on Free anyway |
| ~~`@cf/meta/llama-3.1-8b-instruct`~~ | — | — | HTTP 410, deprecated 2026-05-30 |
**So route by cost shape:** embeddings/reranking are ~300× cheaper per token than the
chat models — treat those as free and the LLM lanes as a daily budget. No Worker needs
deploying either way; Workers AI is a plain REST call.
**Cloudflare token gotcha (cost a wrong diagnosis 2026-07-28).** Modern CF credentials
carry a prefix + CRC32 suffix: `cfk_` user API key · `cfut_` user API token ·
**`cfat_` ACCOUNT-OWNED API token**. An account-owned token **does not verify at
`/user/tokens/verify`** — that endpoint answers `1000 Invalid API Token`, which reads
exactly like a bad credential and is not. Use
`GET /accounts/$CF_ACCOUNT_ID/tokens/verify` instead. Also: `wrangler whoami` prints the
**account ID**; a *member ID* is also 32-hex and is NOT interchangeable.
**z.ai reality check (measured 2026-07-28, NOT the marketing claim):** the account's
`/models` lists `glm-4.5 4.5-air 4.6 4.7 5 5-turbo 5.1 5.2`, but every one of those
returns **HTTP 429 "Insufficient balance or no resource package"**. Only
`glm-4.5-flash` (not in the list) actually answers. The widely-repeated "300M free
tokens/day for GLM-5.2" did **not** hold on a real key — GLM-5.2 needs a paid balance.
- **Reasoning models return an EMPTY answer on too small a budget** — success-shaped
failure. `glm-4.5-flash` spends the budget on a hidden `reasoning_content` channel
before writing `content`: at `--max-tokens 120` it returned HTTP 200 with **zero**
content, at 3000 a full correct answer. `gemini.py` now logs a `WARNING: empty content`
line naming the cause. Give any reasoning model ≥1500. (Same shape as the gpt-oss:20b
8192-truncation lesson — reasoning overhead is not free budget.)
- **Witness drift sensor (trust has a forgetting curve):** local runs are silently re-run
on `SMITH_WITNESS_MODEL` (default gpt-oss:20b) on an FSRS-style schedule — the interval
grows with each consecutive agreement (every 4 runs -> 8 -> 16 ... cap 256) and COLLAPSES
back to 4 on any disagreement or verified-bad verdict. Verification never reaches zero;
stable routes just earn longer intervals. Disagreement = DRIFT SIGNAL in the report
(never an auto-verdict). `SMITH_WITNESS_RATE` switches back to flat-rate sampling.
**Comparator fixed 2026-07-18:** it was exact-match on raw text, applied to one-shot CODE
generation — a markdown fence or a type hint counted as drift, so the sensor sat at 18%
agreement, every model's streak stayed pinned at 0, the interval never grew, and the
logged DRIFT entries before this date are mostly spurious (ignore them). Now: fences are
stripped, and when both outputs parse as Python they're compared as AST skeletons with
annotations/docstrings/comments normalized away — so formatting differences agree while
a real logic change, a different algorithm, or a differing classification still flags.
## Troubleshooting
- `GEMINI_API_KEY is not set` → user exports it, fresh shell.
- HTTP 429 → auto-retries; persistent → switch flash↔pro or `gemini-cli`.
- Model not found → `--list-models`. Weak output → tighten prompt, lower temp, escalate pro.
- Blocked/empty → block reason is on stderr; rephrase.