augustus · v0.3.0 · 2026-09-20 · sha256 2e0b4bfae10d50dc

augustus v0.3.0A

Immutable. This exact content is served forever at /api/v1/blob/2e0b4bfae10d50dc.

---
name: augustus
description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"evidence-preserving stdout prune (not summarize)\", \"hard token/format envelope then soft Noul\", \"fail-safe keep original on prune failure\", \"stdout prune vs session compaction\", \"specialist S1 computer-use (Cua-S1 form-v0; not TypeSafe Jev)\", \"plan ≠ execute / dry-run default\", \"observed-element option head (fill/check/click/skip)\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, \"uncalibrated local likelihoods vs Noul / CUDA replica\", \"decision-native RAG retrieve wide then decide then evidence set\", \"classify-first MCP / read selectively\", \"living applied-mappings atlas / class patterns not a 342 hit list\", \"draft-gate silence as safer / heartbeat\", \"robotics text-state vs pixels\", \"verbatim session ledger / scored recall\", \"judgment as language primitive / English-as-config\", \"pre-registered AMBIGUOUS eval / cascade sign-flip\", \"healthcare Harbor-shaped S1+S2\", \"pre-exec tool gate allow/block/review\", \"productized public primitive / judgment wall\", \"meaning-search without embeddings\", \"attention≠correctness PR review\", \"skills→oxlint / AST prove ∩ remainder\", \"session-sticky first-prompt routing\", \"measured RAG rerank vs generative rerank\", \"Stagehand extract pick-and-copy / judge\", \"harness observe-score-act productization\", or \"formally verify with Jev\", \"capability kernel / secrets never in the agent\", \"Jev is SENSOR not policy\", \"type-safe ≠ correct\", \"typed control plane around DSPy\", \"native-probability calibration / Brier/ECE arena\", \"fan-out as measurement economics\", \"engine owns truth / Jev owns judgment\", \"human-confirmed kill gate\", \"train specialist when downstream reads p vs few-shot hosted when only argmax\", \"decide→policy→LLM leftover cascade\", \"Noul 0.5 cannot-tell never rounded\", \"calibration ≠ sortable / ORDER BY over Jev probs\", \"pairwise inversion / Score ordinality / two-decimal ties\", \"wire-compat self-hosted /v1/systemone GLiFormer\", \"class-backend economics\", \"loopback gateway hosted + local OpenJev\", \"do not distill Jev as teacher of record\", \"active-learning triage / training-data VOI\", \"index-once ask-many / citable evidence packets\", \"meaning-grep AND/OR/NOT line Nouls\", \"closed-vote-only computer-use / no planner LLM\", \"Jev vs local MLX PCD Harbor\", \"PCD O(1) speed ≠ calibrated Noul\", \"host-owned handlers × System One\", \"OMP/pi fail-open acceptance gate\", \"permission vs probability / operator owns the safety bar\", \"judgment ≠ permission / Jev never grants access\", \"eval integrity / instrument not score / dinostomp jev-as-if\", \"constrained optimizer + S1 features / never sole hot-path gate\", \"privilege ≠ verdict / effect contracts not tokens\", \"attention filter / VOI for human review / never blocks / never green unless sure\", \"measurement owns endorsement / evidence-gated question packs\", \"Jev supplies evidence / code owns authority\", \"ranking ≠ calibration / never hard-threshold raw p as frequency\", \"hot-click CU / indexed element table / S1 on click path\", \"Jev judges relevance / code decides structure / never rewrite\", \"local rules first then remainder / never auto-train on model's own hides\", \"combinators / System One as control plane / not chat turns\", \"receipts not leaderboard / type-safe ≠ correct jaggedness\", \"VOI over skill library / skillranker abstention\", \"OOD calibration / AUC ≠ ECE / sign of miscalibration by type\", \"Jev vs thinking-budget small models / frontier-100\", \"turnstile / replayable evidence≠authority\", \"MLX one-pass schema→JSON / Apple Silicon replica economics\", \"memory leases ended by new evidence\", \"never confidently wrong / TLA+ compose with judgment / escalate instead of hard-gate\", \"no seal no advance / coverage ledger / mint ≠ product brain\", \"skill-broker sibling turnstile/skillranker / judgment ≠ permission\", \"sureness / CERTAIN|CONFIDENT|LEANING|TORN|CLUELESS / max_prob is generous\", \"JevBench Harbor/jevals practice / calibration not in Main Score\", \"CI typed gate before expensive review / ci-gatekeeper\", \"Codex MCP host adapter / jev_select_capability\", \"judgment as attention redirect not merge blocker / jev-preflight\", \"compress-before-first-send / dizk jev-lens vs rashed attention filter\", \"tools≠use / SessionStart over hoping the model recalls\", \"observational memory / keep-kind verbatim / pi-om\", \"open-Jev class / openvons / JevPick menu decode\", \"physical-world System One / HA-Jev / not for locks\", \"judgment outside the store / jevql CLI\", \"landed-script trust / headless≠auto-approve\", \"digital-design combinators / extended Router Loop Retry Fallback Memory\", \"VOI cache admission / same-intent skip LLM\", \"BM25 vs Jev skill routing Harbor harness\", \"zeroshot vs BERT / contamination DiD / label-equivalence\", \"typed escalate continue abort baton / inverted loop\", \"worth-your-attention VOI / ThinkyMiner Winnow vs kevinpita winnow\", \"Jev WHETHER Python HOW LLM WHAT\", \"conflict vs ignorance / named Choice escape\", \"Playwright executes Jev chooses / sample-from-distribution\", \"OpenJev /v1/decide not TypeSafe drop-in\", \"SemIf wire-compat runoff\", \"decision-as-memory flywheel\", \"record/replay CI / jevassert\", \"failure-finding arena / jevarena ≠ jev-arena\", \"BBQ stereotype/uncertainty/cost\", \"decider≠executor / jeffrey\", \"sentence-as-rule lint / jevlint ≠ JevLint\", \"VOI hunk prune / prune-review\", \"whole-repo intent VERIFIED/VIOLATION/UNKNOWN\", \"GLiNER2 System One spec ≠ replica\", \"Rust/WebGPU grande / Clojure Laya byte parity / CPU SemIf\", \"ONNX ModernBERT local-jev measured not equivalent\", \"persist constraints across compaction / pi-heed\", \"calibration+cost as first-class gates\", \"Harbor-shaped Jev vs schema-guided LLM-as-judge / jev-judge-bench ≠ jevarena ≠ jevbench\", \"hand no-text steps to Jev / jev-use / Vercel drops confidence / margin fallback\", \"Pi System-One control plane / pi-jev-control\", \"generation as tree of Choices / never free-generates / jev-gpt\", \"OpenRouter recipe atlas / samples not benches / jev-cookbook\", \"personal history feed / no social graph / jevfeed\", \"competing NAR claims / dual-channel ECE / claim-verification / openJev-verdict ≠ OpenJev\", \"empty compaction-proxy skip / IPECTER\", \"throughput ≠ latency / like-for-like ECE\", \"1-token logprob endpoint ≠ Noul / coverage ≠ correctness / chakuho\", \"open replica engine / jevinf / argmax-parity ≠ ECE\", \"unofficial Elixir SDK ≠ OTP peer / dannote/jev\", \"jevex rename + n=16 SWE VOI / files-to-read\", \"commit pre-review attention≠verdict / middle band never rounded / commitjev\", \"Hermes plugin is Agnes not TypeSafe\", \"pi-jev-compact ≠ pi-jev-compaction / verbatim summarizer replacement\", \"empty Codex-proxy skip / IPECTER runway\", \"decision-native inbox / mailordinal / humans own ambiguity\", \"unofficial jev-cli not ready / ≠ jevql\", \"laya-multilingual / English checkpoint confident-wrong OOD / ships uncalibrated\", \"schema-conditioned DeBERTa scorer / peaked ranking ≠ calibration\", \"HF 401 access / GitHub 404 Hub-only\", \"productized System One HTTP / classifier.dev / label+confidence public contract\", \"escalate-under-threshold / smart tier 0.7 / multi-label ignores tier\", \"silent-fallback FALLBACK marker / granite 0.546 vs advertised 0.800\", \"vs_jev tracked JSON not transcription / read eval/README before quoting\", \"choxos/jev-reviewer ≠ egma-ai / systematic-review pointer-not-generator\", \"two-pass Choice+Noul / relative which-line + absolute does-this-line\", \"not-found is an answer / no paraphrase invent\", \"human check as productized judgment / checked never overwritten\", \"githubnext/localjev ≠ kunchenguid/local-jev / prompted JSON ≠ structured logit read\", \"wire-compat ≠ logit-equiv / self-reported probs / entropy confidence\", \"institutional open-replica / GitHub Next /v1/systemone\", \"Harbor-shaped bake-off AG News BoolQ SST-5 / 1200-request caveats\", \"LM Studio runner gap / structured-read primitives for OpenJev parity\", \"NandhaKishorM/laya packaging ≠ Hub-only / Router script-before-p\", \"post-T ECE ≠ raw ECE / Banking77 token-budget / 0.85 still soft / not TypeSafe drop-in / external census ≠ scored bake-off / GLiNER2+routers class-boundary / incomplete vs watch / Harbor honesty watch / JevBench v1.2 geometric-mean I/C/S/K / cal now ON rank / weight sensitivity / option-order 72→21 / instruction models class-boundary / ×2 latency assumption / est. costs / Laya absent gap / Qwen3.8 27B ≠ Archer\", \"hourly already-folded watch / apply-the-five / skip thin noise\", \"hard-gate Noul as PR gate is soundness theater / totally-tim/jev-gate ≠ jev-gateway\", \"S1 never stalls waiting / S2 one-use advisory\", \"purple telemetry = consumed not arrived\", \"Local controller ≠ githubnext/localjev\", \"seed = geometry not async replay\", \"20% starting gate still soft / schema-safe ≠ correct\", \"no pixels to either provider / confidence ≠ selected probability\", \"experimental viz not a flight controller / S2 never grants\", \"OCR+AX observe-score-act / typesafe-computer-use\", \"never send screenshot to frontier for the decision\", \"overlapping CU options = false low confidence\", \"split kind/item/site / offscreen\", \"writer/decider split + post-type Noul still soft\", \"155× one-screenshot Harbor-shaped ≠ taskset\", \"AX never sole / Spotify 0\", \"decision ≠ answer-reader capture\", \"typesafe-computer-use ≠ jev-ultrafast ≠ cua-s1 ≠ camoufox\", \"ASR observe-score-act / jev-voice-browser\", \"partial-speech VOI / complete Noul / free-text waits\", \"spoken confirm ≠ hard auth\", \"numbered overlay disambiguate without another model\", \"moritzkremb/jev-voice-browser ≠ jev-voice-control ≠ nikolas-j\", \"wrap-as-execution / AgentGhost ALLOW ASK DENY\", \"rules first then Jev remainder / ASK throws / fail-closed\", \"reddpy/AgentGhost ≠ jwen5419807/agentghost ≠ vventirozos\", \"JP genre atlas / studio_yebisu / stars ephemeral ≠ eval\", \"Jev Clearly Explained / akshay_pachaar / LLM hammer\", \"schema-safe ≠ correct / 200× 400× TypeSafe ceiling\", \"questions-as-code / shadow first / not a TypeSafe how-to\", \"proposition ≠ embedding / contrast-set refund\", \"boolean composition of soft Nouls / AND OR NOT after threshold\", \"uehaj/jev-semgrep ≠ semgrep.dev\", \"meaning-grep dedicated fold / not a gate\", \"decision-validated UI / Jev never authors text / gram-render\", \"decision-as-assert / jevtest ambiguous band\", \"typed decisions drive UI / jev2ui\", \"hybrid S1 closed verb menu / anima3 / jeff confidently flat\", \"pointer-not-generator search / JevFind\", \"jev-frontier-bench ≠ frontier-100 / ChaosNLI JS\", \"product bakeoff ≠ architecture duel / jev-gliclass-bench\", \"four engines same questions / majority floor / calibration ≠ discrimination\", \"authorship named escape / not courtroom evidence\", \"ha-switchboard HA remains execution / ≠ HA-Jev\", \"n8n classify/route/score / Low Confidence abstention\", \"fast-jev-compaction-pi ≠ pi-jev-compact ≠ pi-jev-compaction\", \"jevloop full-distribution optimizer / no LLM in the loop\", \"laya-vision SmolVLM / score untrained / ≠ blackwood ≠ Archer\", \"Cerebellum-2B /v1/decide ≠ TypeSafe / wire-compat vs agent-routing\", \"laya-grounded not drop-in / phishing regress / Platt not temperature\", \"GestaltLabs/Jeff-1 ≠ logan-markewich/jeff / acc vs ECE n=9730\", \"stanley-code empty findings ≠ approval / human promote\", \"findme ≠ JevFind / NL memory beam-search FS\", \"jevsubrouter price workers not conversation / counts ≠ dollars\", \"feelings .feels() default 0.5 is Noul-0.5-never-rounded / ≠ hunch ≠ Probably\", \"apa-agent-harness ≠ AntonioCoppe/jev-harness / unpublished npm\", \"grok-bot-jev skill cannot force a bot that ignores it / A/B proxies not tokens\", \"Essentiel-Jev never authority / human every action\", \"enzo-mcp independently falsifiable claims / ≠ jev-sift\", \"pigeonhole OTHER skip / decision-as-filing\", \"jev-reliability Nothing about accuracy\", \"clduab11/jev-test ≠ realZachi/jevtest / Nothing runs yet\", \"jev-rag-benchmark Jev wins is not an assumption\", \"dairui1/jev-lab ≠ BrendanH18/jev-lab\", \"jevmail gmail.readonly / mailjay archive/trash\", \"ZHUBoer/ego-jev reserved __none__\", \"runWorkflow completed ≠ success\", \"jsort scores are relative\", \"Noul not Choice for scale\", \"groundedness-judge-bench native vs schema-guided\", \"implicit_true included in yes\", \"jev_playground 0 promotions\", \"routing-backtest 0.0447%\", \"yuyang2230/jev-agent-skill jev-1.13-free\", \"jev-techstack-classifier stack_config.json\", \"s1_ruby collapse late\", \"undecided? abstain\", \"2389-research/judgement license null\", \"confidence ≠ winner p\", \"typesafeai-sdk-community not a new species\", \"tpellet/hunch exit 3\", \"never-execute list\", \"jev-file-search scores not calibrated accuracy\", \"jev-linkmap Jev never sees S2 prose\", \"muhammedilyasy/jev-mail metadata only\", \"tidy none-of-folders stay\", \"tab-bouncer pinned/audio/current never closed\", \"lkclean Show fail-open\", \"jev-yt-time-saver Show anyway\", \"ORIGIN pause-if-no-Jev\", \"validResponse sums-to-1\", \"jev-crawlers risk bands never raw boolean\", \"jevbrain AUTO_ACT is not a Noul\", \"judgekit YAML classify/score/route/verify\", \"typed-judge-kit verdict-in-code\", \"alsoleg89/decide packing VOI\", \"0.8 ≠ 80% accuracy\", \"Jev-Calibration Platt ECE 0.117→0.052\", \"jev-calibration-arena never acts\", \"ctmx/openrouter-jev-mcp Decision-as-Plugin\", \"FrancoisChastel/jev-code ≠ npm jev-code\", \"claudecode-jev-marketplace fail-open not hot path\", \"pedroknigge/mcp_jev packs not ask_jev\", \"cyrusasco/typesafe-mcp noul deadband 0.35–0.65\", \"codaaiteam/jev-skill jevtypesafeai.com ≠ TypeSafe\", \"hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev\", \"nanoprune 2.8MB ECE 2.58%\", \"smartdio/jev-browser-agent ≠ ZHUBoer/ego-jev\", \"Dakai/omp-jev-web DONE ≠ proof\", \"hari007sh/jev ≠ dannote/jev\", \"0thernet/system-one-skills deterministic verify\", \"typed-gate band [0.40,0.60] is refusal\", \"pi-jev-gate fail-closed; choice is the verdict\", \"Foq ~25ms/2.2GB local\", \"rev prefill-only + HF jev-0.5b\", \"robfrase/jev planning memo\", \"typesafe_agent_gates 27/27 / 31/31\", \"EpicEric/safe-sh static remainder\", \"pastepilot Confirm before act\", \"Jev-Reranker live Jev not yet measured\", \"sessionwise opt-in relevance\", \"jev-search pointer sieve\", \"400ms Salesforce WebMCP\", \"typesafe-scheduler-diagnostics advisory\", \"droidjev screenshot-free\", \"Tewoto1 jevcu planner still writes\", \"ha-conversation-jev Jev→Grok\", \"dsh-jev can only gate\", \"jev-classification-benchmark specified not run\", \"jev-luna-pagerduty p≥0.50\", \"meldltd/meldecision laya-go ONNX\", \"laya-doom never pixels\", \"logixism/laya-api empty README\", \"akpsahan/laya ≠ Archer\", \"choxos/jevchess engine owns truth\", \"jev-drive sim not AV\", \"story-arc Jev never authors\", \"jev-hs-assistant HS6\", \"golergka/jev-plays-starcraft-2 UI-verified ≠ API Victory\", \"awesome-jev-use-cases catalog\", \"Nibir1/typesafe-go ≠ official\", \"fingerprint after redact\", \"recall vs decide\", \"publish fingerprints+answers\", \"CI replay as Harbor cousin\", \"Cache hit ≠ correctness\", \"hyperspaceai/jevcache ≠ kushals256/jevcache\", \"human labels only\", \"score never auto-accepts\", \"production capture flywheel\", \"sutro-sh/jev-align ≠ caiovicentino/jev-align\", \"guidance ≠ hook\", \"catalysts ≠ summaries\", \"compile-time System One\", \"unofficial ≠ TypeSafe\", \"format_version modernbert-jev/1\", \"Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev\", \"LFM default ≠ ModernBERT backend\", \"Nemotron ≠ TypeSafe Jev\", \"not a calibrated replacement\", \"djev-dev complements djev-spark\", \"images as Choice options\", \"Laya essay numbers *theirs*\", \"Router/OOD confidence\", \"hosted bootstrap ≠ silent TypeSafe\", \"difficulty + policy thresholds + JSONL trace\", \"jev-codex-pilot model + reasoning depth\", \"keep/shadow/hybrid/reject\", \"quarry evidence projection\", \"Frank-ZY-Dou/awesome-jev robotics/3D/control\", \"one-dollar-tahoe TypeSafe Jev defense eval\", \"jevguard calibrator/cache/escape\", \"jev-ci-selector CI shadow mode\", \"llama-jev llama.cpp replica\", \"petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator\", \"seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard\", \"webNeat/llama-jev ≠ WiktorB2004/llama-index-jev\", \"OpenCode jev-pruner context sieve\", \"observe→score-candidates→prune\", \"jev-zen / jev-1.13-free\", \"zen-chat ≠ Noul\", \"fail-open original\", \"keepScore >0.1 floor\", \"host port of tamaratran/jev-pruner\", \"indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode\", \"jev-webagent-bench empty stub\", \"Kiln-AI/jev_jsonschema noul_threshold 0.5\", \"NSStudent/JevSwiftSDK unofficial\", \"GLiNER2 native Apple path\", \"unofficial Swift/Core ML GLiNER 2.5-small\", \"entity spans + confidence\", \"not Choice/Score/Noul\", \"not TypeSafe\", \"label descriptions as schema\", \"on-device ANE economics\", \"honesty locks\", \"shershah1024/gliner-native-runtime ≠ Fastino\", \"≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx\", \"default threshold 0.1 still soft\", \"Decision Graph Protocol frame→assess→commit\", \"app retains permissions/effects\", \"Jev-first assessor-neutral\", \"guarded commit / receipt/next frame\", \"assessment batching\", \"hard-gating DGP as safety theater\", \"numerous-com/dgp ≠ TypeSafe official\", \"jegrep calibrated path+range Nouls\", \"no embeddings/index/daemon\", \"~$0.01–0.03 typical\", \"agent --json\", \"can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep\", \"Archer-arch fidelity\", \"kev family OOD 0.76–0.77 vs Jev 0.86\", \"block-causal isolation\", \"pointer/readout CE-trained\", \"/v1/systemone drop-in\", \"replica honesty\", \"cost-sensitive decision theory × System One probabilities → control flow\", \"thresholds derived from costs not hard-coded\", \"YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human\", \"auto-batching same-object questions\", \"Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch\", \"judgment vs generation\", \"deterministic execution after probabilistic judgment\", \"exactly one app-owned callback\", \"explicit uncertain branch\", \"Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit\", \"variable-N option scoring as the trainable object\", \"dynamic candidate bags not fixed label sets\", \"zwliJay/jev-forge ≠ NanoJev\", \"open replica economics / latency vs closed Jev\", \"NAR local drop-in\", \"wfzyx/von late-catch HIGH\", \"competing NAR claims / replica honesty\", \"typed judgments vs chat judges on guardrailing\", \"ishaannk/llm-vs-jev cross-note only\", \"deeper integrity fold is rh-guard\", \"nothing wins outright\", \"can be argued out of guarding\", \"Jev IS the if-statement\", \"judgments/probabilities drive branches\", \"text model only writes prose\", \"interpreter owns variables/loops/budgets/replay\", \"otherwise maybe / confidence gate\", \"chaos samples after the gate\", \"southpolesteve/probably ≠ carldaws/hunch ≠ feelings ≠ Kungie/gut ≠ Illusion47586/judge ≠ tidymodels/probably\", \"133★ / forks 10 live\", \"build calibrated classifiers from human feedback\", \"retrieve by relevance not resemblance\", \"one calibrated yes/no per memory in one request\", \"pointer mode 17/18 19/20 *theirs*\", \"embedding resemblance misses the allergy\", \"samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate\", \"memory leases ended by new evidence\", \"six Nouls then fixed rules in code\", \"0 of 157 false invalidations\", \"questions/plans/directives are not evidence\", \"unsure → review queue\", \"host keeps the store\", \"name↔body / comment truth / test-claims\", \"mizchi/jev-lint is mizchi/jevlint rename\", \"no shipped rule has severity error\", \"~1 in 5 findings wrong *theirs*\", \"mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint\", \"JSON Schema → typed JSON via Jev\", \"noul_threshold 0.5 decoder not a proof\", \"IncompatibleSchemaError lists every bad property\", \"on-device Laya CoreML ANE\", \"~5 ms P50 short decisions\", \"189/189 FP16 checkpoint parity\", \"10× not achieved\", \"mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya\", \"softmax over allowed tokens ≠ Noul\", \"question-first cache\", \"Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge\", \"Jev-first Pi agent loop\", \"slow-LLM fallback\", \"explicit action menu / CandidateSource unimplemented\", \"62 tests wiring not quality\", \"direwolfiy/JevPi ≠ standardagents/jevpilot ≠ pi-jev-control\", \"resume-screening bias audit methodology\", \"name×resume factorial independent Nouls\", \"callback determined by resume quality\", \"mean-probability name gaps operationally negligible\", \"natemoo-re/bias-bench ≠ BBQ\", \"Plan/PRD panel → code-owned pass|review|block\", \"cheerleading out of scope\", \"austindixson/planalyzer ≠ single-goodness Noul\", \"cost-aware multi-model routing/escalation\", \"decide vs do\", \"successful-task cost\", \"cannacre8ive/switchboard-ai ≠ ha-switchboard ≠ hermes-switchyard\", \"frozen-protocol zero-shot bench\", \"TypeSafe Jev vs PrismNLI vs Laya\", \"contamination caveat\", \"elcronos/jev-vs-open-decision-models ≠ JevBench ≠ DMB\", \"context-window admission control\", \"VOI gate which tokens are worth the expensive model\", \"fail polarity per lens\", \"on small inputs lenses lose money\", \"cvsgireesh/jevusher ≠ jev-sift ≠ winnow\", \"typed decision control plane\", \"receipt ≠ authorization\", \"historical-v0 zero retained cases\", \"MokiMeow/jev-fabric ≠ jev-forge ≠ dgp\", \"live 15-dim typed rubric re-score per pause\", \"scoring economics exemplar\", \"OpenJev/Codiv ≠ TypeSafe hosted\", \"jose-troche/live-rubric ~$0.000004 desc / ~$0.000006 README\", \"adversarial pre-registered Jev eval\", \"28 predictions before data\", \"123,805 requests\", \"confidence does not track ignorance\", \"polite injection 65% / crude 0%\", \"willkelly/jev-evaluation ≠ jevals ≠ jev-baselines-eval\", \"provider-neutral Elixir/BEAM Noul/Choice/Score SDK\", \"class infrastructure\", \"nshkrdotcom/system_one_sdk ≠ typesafe_sdk ≠ dannote/jev\", \"soft Noul ≠ hard safety\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)."
license: MIT
metadata:
  version: 0.3.0
  typesafe_skill: v0.5.7
  typesafe_skill_commit: 65a39f3
  tribute: "Named for Augustus De Morgan (1806-1871), mentor of William Stanley Jevons."
---

# Augustus

Design-judgment skill for **placing typed probabilistic judgment** (the
Jev-class of System One / decision models) using mathematical, logical, and
algorithmic mental models — across **AI, software, business, knowledge
work, and life**. Not limited to software engineering. Formal methods
are one pillar (`references/formal-methods.md`); the portable frames
are `references/mental-models.md`.

TypeSafe Jev is the documented **exemplar** (typed Choice / Score /
Noul), not the monopoly. This skill owns **where judgment belongs**;
the official `typesafe-ai` skill plus the live docs own Jev integration
contracts — read them before writing Jev API code. Neighbor skills
`tenbin` (lint/measure) and `decision-first` (try-Jev-first habit) own
their jobs. Do not collapse into a TypeSafe how-to, a Laya install, a kev
serve, a blackwood vLLM how-to, a jev-local Docker install, or a GLiClass or GLiNER tutorial.

Pick the **pillar** from the hole (expected utility, VOI, MCDA, signal
detection, search/control, org/safety, formal methods), then the
**family** (`references/judgment-class.md`), then the vendor. Default
placement is **mixed architecture**: exact work in code/policy, narrow
judgment on a System One–class model, generation only where something
must be written.

Central model: **evidence → semantic judgments → explicit policy → checked
action → observed outcome.** Every design must name what the judgment
model estimates, what remains exact, what (if anything) is still
generated, and what experiment could prove the idea wrong.

For genuinely new problem shapes, use the toolbox sweep
(`references/toolbox-mapping.md`): find the judgment-shaped component of a
classical method you already trust, substitute it, classify the win
(marginal / newly-feasible / invalid), and falsify.

## Protocol

1. If the request is "replace the LLM/stack with Jev", "isn't this just
   classification?", "is Jev probabilistic programming?" (marginals vs
   joint — not a PPL), "low vs high entropy / frontier model for
   this?", "is Jev the only model?", "is this only for
   software?", "formally verify with Jev / replace TLA+ / Dafny /
   DST", "Alloy vs Apalache", "GLiNER vs Jev", "LLM-as-judge",
   "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul",
   "Jev inside the database / sqlite-jev", "Jev picks bitrate / join
   order / the model", "wait for Archer", "lint the request / missing
   other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "OpenCode jev-pruner vs Claude jev-pruner", "zen-chat vs jev-zen Noul", "hard envelope then Noul prune", "Cua-S1 vs TypeSafe Jev", "plan vs execute dry-run", "specialist computer-use vs general agent", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", "combinatorial grid vs extractive", "uncalibrated local likelihoods", "decision-native RAG", "classify-first / read selectively", "living applied-mappings atlas / class patterns", "silence as safer / draft-gate heartbeat", "robotics text-state vs pixels", "verbatim ledger vs summary", "judgment as language primitive", "Stagehand extract pick-and-copy", "harness observe-score-act vs demo loop", "public judgment wall / six parallel questions", "meaning-search without embeddings", "attention ≠ correctness", "skills→oxlint / AST prove ∩ remainder", "session-sticky first-prompt routing", "measured RAG rerank vs generative rerank", "capability kernel / secrets never in the agent", "Jev is SENSOR not policy", "type-safe ≠ correct", "typed control plane around DSPy", "native vs verbalized confidence", "engine owns truth / Jev owns judgment", "human-confirmed kill gate", "train specialist vs few-shot hosted", "decide→policy→LLM leftover", "Noul 0.5 cannot-tell never rounded", "calibration ≠ sortable / ORDER BY", "pairwise inversion / Score ordinality / two-decimal ties", "wire-compat GLiFormer /v1/systemone", "class-backend economics", "loopback gateway hosted + local", "do not distill Jev as teacher", "active-learning triage", "evidence-packet explorer", "meaning-grep AND/OR/NOT", "closed-vote-only / no planner LLM", "Jev vs PCD Harbor", "PCD O(1) ≠ Noul", "host-owned handlers × System One", "OMP/pi fail-open gate", "permission vs probability / operator owns thresholds", "judgment ≠ permission / Jev never grants access", "eval integrity / instrument not score", "constrained optimizer + S1 features / never sole hot-path gate", "privilege ≠ verdict / effect contracts not tokens", "attention filter / VOI for human review / never blocks / never green unless sure", "measurement owns endorsement / evidence-gated question packs", "Jev supplies evidence / code owns authority", "ranking ≠ calibration / never hard-threshold raw p as frequency", "hot-click CU / indexed element table", "Jev judges relevance / code decides structure", "local rules first then remainder / never auto-train on own hides", "combinators / System One as control plane", "receipts not leaderboard / type-safe ≠ correct jaggedness", "VOI over skill library / skillranker abstention", "OOD calibration / AUC ≠ ECE", "Jev vs thinking-budget small models", "turnstile / replayable evidence≠authority", "MLX one-pass schema→JSON / Apple Silicon replica economics", "memory leases ended by new evidence", "never confidently wrong / TLA+ compose / escalate instead of hard-gate", "no seal no advance / coverage ledger / mint ≠ product brain", "skill-broker sibling / judgment ≠ permission", "sureness bands / max_prob is generous", "JevBench / calibration not in Main Score", "CI typed gate before expensive review", "Codex MCP host adapter", "judgment as attention redirect / jev-preflight", "compress-before-first-send / dizk jev-lens", "tools≠use / SessionStart over hoping", "observational memory / pi-om keep-kind", "open-Jev class / openvons / JevPick", "physical-world System One / HA-Jev / not for locks", "judgment outside the store / jevql", "landed-script trust / headless≠auto-approve", "digital-design combinators / extended five", "VOI cache admission / same-intent skip LLM", "BM25 vs Jev skill routing Harbor harness", "zeroshot vs BERT / contamination DiD", "typed escalate continue abort baton / inverted loop", "worth-your-attention VOI / ThinkyMiner Winnow", "Jev WHETHER Python HOW LLM WHAT", "conflict vs ignorance / named Choice escape", "Playwright executes Jev chooses", "OpenJev /v1/decide not drop-in", "SemIf wire-compat runoff", "decision-as-memory flywheel", "record/replay CI / jevassert", "failure-finding arena / jevarena ≠ jev-arena", "BBQ not a bias cert", "decider≠executor", "sentence-as-rule lint / jevlint", "sentence-as-rule lint / jev-lint is jevlint rename", "VOI hunk prune", "whole-repo intent VERIFIED/VIOLATION/UNKNOWN", "GLiNER2 spec ≠ replica", "open replica substrates / grande / laya-jolt / JEV-CPU", "ONNX local-jev not equivalent", "persist constraints across compaction / pi-heed", "calibration+cost first-class gates", "Harbor-shaped Jev vs SGR LLM-as-judge / jev-judge-bench ≠ jevarena ≠ jevbench", "hand no-text steps / jev-use / Vercel drops confidence", "Pi System-One control plane / pi-jev-control", "never free-generates / jev-gpt tree of Choices", "OpenRouter recipe atlas / samples not benches", "personal history feed / jevfeed / no social graph", "competing NAR claims / dual-channel ECE / openJev-verdict ≠ OpenJev", "empty compaction-proxy skip / IPECTER", "throughput ≠ latency / like-for-like ECE", "1-token logprob endpoint ≠ Noul / coverage ≠ correctness", "open replica engine / jevinf", "unofficial Elixir SDK ≠ OTP peer", "jevex n=16 files-to-read VOI", "commit pre-review attention≠verdict / middle band", "Hermes plugin is Agnes not TypeSafe", "pi-jev-compact ≠ pi-jev-compaction", "empty Codex-proxy skip / IPECTER runway", "decision-native inbox / mailordinal", "unofficial jev-cli not ready / ≠ jevql", "laya-multilingual / English checkpoint confident-wrong OOD", "schema-scorer peaked ranking ≠ calibration", "HF 401 / GitHub 404 Hub-only", "productized System One HTTP / classifier.dev", "escalate-under-threshold / smart tier / multi-label ignores", "silent FALLBACK / granite 0.546 vs advertised 0.800", "vs_jev tracked JSON / read eval/README", "choxos/jev-reviewer ≠ egma-ai / systematic-review pointer", "two-pass Choice+Noul evidence extraction", "not-found is an answer", "human check as productized judgment", "githubnext/localjev ≠ kunchenguid/local-jev", "wire-compat ≠ logit-equiv / prompted JSON ≠ structured read", "self-reported probs / entropy confidence", "GitHub Next local /v1/systemone", "LM Studio runner gap / structured-read primitives", "NandhaKishorM/laya packaging ≠ Hub-only / Router script-before-p", "post-T ECE ≠ raw ECE / Banking77 token-budget", "0.85 still soft / not TypeSafe drop-in", "external census ≠ scored bake-off", "GLiNER2+routers class-boundary", "incomplete openjev census vs watch", "Harbor honesty watch / silent fallback", "JevBench v1.2 geometric mean / cal ON rank / weight sensitivity", "option-order 72→21 / instruction models in the class table", "self-host latency ×2 assumption / est. costs", "Laya absent is a gap not a named exclusion", "Qwen3.8 27B ≠ Archer", "hourly already-folded watch / apply-the-five / skip thin noise", "hard-gate Noul as PR/quality gate is soundness theater", "S1 never stalls waiting / S2 one-use advisory", "Local controller ≠ githubnext/localjev", "purple telemetry = consumed not arrived", "seed = geometry not async replay", "20% starting gate still soft", "no pixels to either provider", "OCR+AX observe-score-act / typesafe-computer-use", "never send screenshot to frontier for the decision", "overlapping CU options = false low confidence", "split kind/item/site", "155× one-screenshot ≠ Harbor taskset", "decision ≠ answer-reader capture", "ASR observe-score-act / jev-voice-browser", "partial-speech VOI / free-text waits", "spoken confirm ≠ hard auth", "numbered overlay without another model", "wrap-as-execution / AgentGhost ALLOW ASK DENY", "rules first then Jev remainder / ASK throws / fail-closed", "reddpy/AgentGhost ≠ jwen5419807/agentghost ≠ vventirozos", "JP genre atlas / studio_yebisu / stars ephemeral ≠ eval", "Jev Clearly Explained / akshay_pachaar / LLM hammer", "schema-safe ≠ correct / 200× 400× TypeSafe ceiling", "questions-as-code / shadow first / not a TypeSafe how-to", "proposition ≠ embedding / contrast-set", "boolean composition of soft Nouls / AND OR NOT", "uehaj/jev-semgrep ≠ semgrep.dev", "meaning-grep dedicated fold / not a gate", "decision-validated UI / Jev never authors text", "decision-as-assert / jevtest ambiguous band", "typed decisions drive UI / jev2ui", "hybrid S1 closed verb menu / anima3", "pointer-not-generator search / JevFind", "jev-frontier-bench ≠ frontier-100", "product bakeoff ≠ architecture duel / GLiClass", "four engines same questions / majority floor", "authorship named escape / not evidence", "ha-switchboard HA remains execution", "n8n classify/route/score / Low Confidence", "fast-jev-compaction-pi ≠ pi-jev-compact ≠ pi-jev-compaction", "jevloop full-distribution optimizer / no LLM in the loop", "laya-vision SmolVLM / score untrained", "Cerebellum-2B /v1/decide ≠ TypeSafe / wire-compat vs agent-routing", "laya-grounded not drop-in / Platt not temperature", "GestaltLabs/Jeff-1 ≠ logan-markewich/jeff / acc vs ECE n=9730", "stanley-code empty findings ≠ approval / human promote", "findme ≠ JevFind / NL memory beam-search FS", "jevsubrouter price workers not conversation / counts ≠ dollars", "feelings .feels() default 0.5 is Noul-0.5-never-rounded / ≠ hunch ≠ Probably", "apa-agent-harness ≠ AntonioCoppe/jev-harness / unpublished npm", "grok-bot-jev skill cannot force a bot that ignores it / A/B proxies not tokens", "Essentiel-Jev never authority / human every action", "enzo-mcp independently falsifiable claims / ≠ jev-sift", "pigeonhole OTHER skip / decision-as-filing", "jev-reliability Nothing about accuracy", "clduab11/jev-test ≠ realZachi/jevtest / Nothing runs yet", "jev-rag-benchmark Jev wins is not an assumption", "dairui1/jev-lab ≠ BrendanH18/jev-lab", "jevmail gmail.readonly / mailjay archive/trash", "ZHUBoer/ego-jev reserved __none__", "runWorkflow completed ≠ success", "jsort scores are relative", "Noul not Choice for scale", "groundedness-judge-bench native vs schema-guided", "implicit_true included in yes", "jev_playground 0 promotions", "routing-backtest 0.0447%", "yuyang2230/jev-agent-skill jev-1.13-free", "jev-techstack-classifier stack_config.json", "s1_ruby collapse late", "undecided? abstain", "2389-research/judgement license null", "confidence ≠ winner p", "typesafeai-sdk-community not a new species", "tpellet/hunch exit 3", "never-execute list", "jev-file-search scores not calibrated accuracy", "jev-linkmap Jev never sees S2 prose", "muhammedilyasy/jev-mail metadata only", "tidy none-of-folders stay", "tab-bouncer pinned/audio/current never closed", "lkclean Show fail-open", "jev-yt-time-saver Show anyway", "ORIGIN pause-if-no-Jev", "validResponse sums-to-1", "jev-crawlers risk bands never raw boolean", "jevbrain AUTO_ACT is not a Noul", "judgekit YAML classify/score/route/verify", "typed-judge-kit verdict-in-code", "alsoleg89/decide packing VOI", "0.8 ≠ 80% accuracy", "Jev-Calibration Platt ECE 0.117→0.052", "jev-calibration-arena never acts", "ctmx/openrouter-jev-mcp Decision-as-Plugin", "FrancoisChastel/jev-code ≠ npm jev-code", "claudecode-jev-marketplace fail-open not hot path", "pedroknigge/mcp_jev packs not ask_jev", "cyrusasco/typesafe-mcp noul deadband 0.35–0.65", "codaaiteam/jev-skill jevtypesafeai.com ≠ TypeSafe", "hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev", "nanoprune 2.8MB ECE 2.58%", "smartdio/jev-browser-agent ≠ ZHUBoer/ego-jev", "Dakai/omp-jev-web DONE ≠ proof", "hari007sh/jev ≠ dannote/jev", "0thernet/system-one-skills deterministic verify", "typed-gate band [0.40,0.60] is refusal", "pi-jev-gate fail-closed; choice is the verdict", "Foq ~25ms/2.2GB local", "rev prefill-only + HF jev-0.5b", "robfrase/jev planning memo", "typesafe_agent_gates 27/27 / 31/31", "EpicEric/safe-sh static remainder", "pastepilot Confirm before act", "Jev-Reranker live Jev not yet measured", "sessionwise opt-in relevance", "jev-search pointer sieve", "400ms Salesforce WebMCP", "typesafe-scheduler-diagnostics advisory", "droidjev screenshot-free", "Tewoto1 jevcu planner still writes", "ha-conversation-jev Jev→Grok", "dsh-jev can only gate", "jev-classification-benchmark specified not run", "jev-luna-pagerduty p≥0.50", "meldltd/meldecision laya-go ONNX", "laya-doom never pixels", "logixism/laya-api empty README", "akpsahan/laya ≠ Archer", "choxos/jevchess engine owns truth", "jev-drive sim not AV", "story-arc Jev never authors", "jev-hs-assistant HS6", "golergka/jev-plays-starcraft-2 UI-verified ≠ API Victory", "awesome-jev-use-cases catalog", "Nibir1/typesafe-go ≠ official", "fingerprint after redact", "recall vs decide", "publish fingerprints+answers", "CI replay as Harbor cousin", "Cache hit ≠ correctness", "hyperspaceai/jevcache ≠ kushals256/jevcache", "human labels only", "score never auto-accepts", "production capture flywheel", "sutro-sh/jev-align ≠ caiovicentino/jev-align", "guidance ≠ hook", "catalysts ≠ summaries", "compile-time System One", "unofficial ≠ TypeSafe", "format_version modernbert-jev/1", "Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev", "LFM default ≠ ModernBERT backend", "Nemotron ≠ TypeSafe Jev", "not a calibrated replacement", "djev-dev complements djev-spark", "images as Choice options", "Laya essay numbers *theirs*", "Router/OOD confidence", "hosted bootstrap ≠ silent TypeSafe", "difficulty + policy thresholds + JSONL trace", "jev-codex-pilot model + reasoning depth", "keep/shadow/hybrid/reject", "quarry evidence projection", "Frank-ZY-Dou/awesome-jev robotics/3D/control", "one-dollar-tahoe TypeSafe Jev defense eval", "jevguard calibrator/cache/escape", "jev-ci-selector CI shadow mode", "llama-jev llama.cpp replica", "petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator", "seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard", "webNeat/llama-jev ≠ WiktorB2004/llama-index-jev", "OpenCode jev-pruner context sieve", "observe→score-candidates→prune", "jev-zen / jev-1.13-free", "zen-chat ≠ Noul", "fail-open original", "keepScore >0.1 floor", "host port of tamaratran/jev-pruner", "indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode", "jev-webagent-bench empty stub", "Kiln-AI/jev_jsonschema noul_threshold 0.5", "NSStudent/JevSwiftSDK unofficial", "GLiNER2 native Apple path", "unofficial Swift/Core ML GLiNER 2.5-small", "entity spans + confidence", "not Choice/Score/Noul", "not TypeSafe", "label descriptions as schema", "on-device ANE economics", "honesty locks", "shershah1024/gliner-native-runtime ≠ Fastino", "≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx", "default threshold 0.1 still soft", "soft Noul ≠ hard safety", "Decision Graph Protocol frame→assess→commit", "app retains permissions/effects", "Jev-first assessor-neutral", "guarded commit / receipt/next frame", "assessment batching", "hard-gating DGP as safety theater", "numerous-com/dgp ≠ TypeSafe official", "jegrep calibrated path+range Nouls", "no embeddings/index/daemon", "~$0.01–0.03 typical", "agent --json", "can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep", "Archer-arch fidelity", "kev family OOD 0.76–0.77 vs Jev 0.86", "block-causal isolation", "pointer/readout CE-trained", "/v1/systemone drop-in", "replica honesty", "cost-sensitive decision theory × System One probabilities → control flow", "thresholds derived from costs not hard-coded", "YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human", "auto-batching same-object questions", "Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch", "judgment vs generation", "deterministic execution after probabilistic judgment", "exactly one app-owned callback", "explicit uncertain branch", "Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit", "variable-N option scoring as the trainable object", "dynamic candidate bags not fixed label sets", "zwliJay/jev-forge ≠ NanoJev", "open replica economics / latency vs closed Jev", "NAR local drop-in", "wfzyx/von late-catch HIGH", "competing NAR claims / replica honesty", "typed judgments vs chat judges on guardrailing", "ishaannk/llm-vs-jev cross-note only", "deeper integrity fold is rh-guard", "nothing wins outright", "can be argued out of guarding"", "Jev IS the if-statement", "judgments/probabilities drive branches", "text model only writes prose", "interpreter owns variables/loops/budgets/replay", "otherwise maybe / confidence gate", "chaos samples after the gate", "southpolesteve/probably ≠ carldaws/hunch ≠ feelings ≠ Kungie/gut ≠ Illusion47586/judge ≠ tidymodels/probably", "133★ / forks 10 live", "build calibrated classifiers from human feedback", "retrieve by relevance not resemblance", "one calibrated yes/no per memory in one request", "pointer mode 17/18 19/20 *theirs*", "embedding resemblance misses the allergy", "samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate", "memory leases ended by new evidence", "six Nouls then fixed rules in code", "0 of 157 false invalidations", "questions/plans/directives are not evidence", "unsure → review queue", "host keeps the store", "name↔body / comment truth / test-claims", "mizchi/jev-lint is mizchi/jevlint rename", "no shipped rule has severity error", "~1 in 5 findings wrong *theirs*", "mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint", "JSON Schema → typed JSON via Jev", "noul_threshold 0.5 decoder not a proof", "IncompatibleSchemaError lists every bad property", "on-device Laya CoreML ANE", "~5 ms P50 short decisions", "189/189 FP16 checkpoint parity", "10× not achieved", "mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya", "softmax over allowed tokens ≠ Noul", "question-first cache", "Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge", "Jev-first Pi agent loop", "slow-LLM fallback", "explicit action menu / CandidateSource unimplemented", "62 tests wiring not quality", "direwolfiy/JevPi ≠ standardagents/jevpilot ≠ pi-jev-control", "resume-screening bias audit methodology", "name×resume factorial independent Nouls", "callback determined by resume quality", "mean-probability name gaps operationally negligible", "natemoo-re/bias-bench ≠ BBQ", "Plan/PRD panel → code-owned pass|review|block", "cheerleading out of scope", "austindixson/planalyzer ≠ single-goodness Noul", "cost-aware multi-model routing/escalation", "decide vs do", "successful-task cost", "cannacre8ive/switchboard-ai ≠ ha-switchboard ≠ hermes-switchyard", "frozen-protocol zero-shot bench", "TypeSafe Jev vs PrismNLI vs Laya", "contamination caveat", "elcronos/jev-vs-open-decision-models ≠ JevBench ≠ DMB", "context-window admission control", "VOI gate which tokens are worth the expensive model", "fail polarity per lens", "on small inputs lenses lose money", "cvsgireesh/jevusher ≠ jev-sift ≠ winnow", "typed decision control plane", "receipt ≠ authorization", "historical-v0 zero retained cases", "MokiMeow/jev-fabric ≠ jev-forge ≠ dgp", "live 15-dim typed rubric re-score per pause", "scoring economics exemplar", "OpenJev/Codiv ≠ TypeSafe hosted", "jose-troche/live-rubric ~$0.000004 desc / ~$0.000006 README", "adversarial pre-registered Jev eval", "28 predictions before data", "123,805 requests", "confidence does not track ignorance", "polite injection 65% / crude 0%", "willkelly/jev-evaluation ≠ jevals ≠ jev-baselines-eval", "provider-neutral Elixir/BEAM Noul/Choice/Score SDK", "class infrastructure", "nshkrdotcom/system_one_sdk ≠ typesafe_sdk ≠ dannote/jev", or "cascade sign-flip / calibration theater":
   read `references/faq.md`,
   then `references/mental-models.md`, then
   `references/mixed-architecture.md`, then
   `references/judgment-class.md` before any mapping. Proof,
   model-checking, contracts, DST, and judgment-vs-proof ownership: also
   read `references/formal-methods.md` (Alloy vs Apalache; DST trio
   Antithesis / Resonate HQ / PufferLib; TOCTOU-of-Noul). One-screen
   alias: `references/formal-semi-formal.md`. Answer with a **placement**
   (sieve / keep-drop / triage / rank / route / gate / perceive /
   abstain / gather / replace-one-classifier-step) and a **pillar +
   family**, not a rewrite, a vendor tutorial, or a Noul-as-proof. If it
   is an existing system, PR, workflow, or practice: also run the
   boundary audit (`references/boundary-audit.md`). Classify each step
   as exact / bounded judgment / generation; recommend the smallest
   insertion, not a redesign. Greenfield with no replacement framing:
   start at step 2.
2. Start from the desired behavior: what the software, person, or
   organization shows, selects, changes, or hands off. Work backward to
   the judgments it needs. Name the domain and the pillar
   (`references/mental-models.md`), then pick the family whose
   *objective* matches the action's fail policy
   (`references/judgment-class.md`). Do not start from a logo.
3. Keep exact work in code, policy, checklists, ledgers, law, and
   recipes: arithmetic, counting, dates, lookups, authorization, safety
   interlocks, control flow, side effects, money. Keep open-ended
   writing, explanation, and code generation on a generative model or a
   person; a judgment-class model may gate, route, or verify around that
   call. Proof, model-checking, contracts, and DST stay with their
   tools — a Noul is a sensor, not a discharged proof obligation
   (`references/formal-methods.md`). A capability kernel keeps
   secrets out of the agent and leaves BLOCK/ASK/ALLOW in policy;
   type-safe is not the same as correct.
4. Give the provider one narrow judgment per question (a knowledgeable
   person could answer in a second given the state). Split multi-factor
   judgments; fuse in code with visible weights. Jev's option/envelope
   limits are Jev's, not the class's — large or changing label sets may
   prefer a GLiClass one-pass (categorize). GLiNER (locate) substitutes
   only where the answer *is* a span in the text
   (`judgment-class.md` species map).
5. Exploit the family's cheap fan-out: batch independent questions in
   one request when the provider supports it (measurement economics:
   200 calibration questions in 2 requests; ~100 Battleship noul/turn); encode text + all labels
   once for GLi\* heads; score many prompts against one
   embedding for dual-encoder vision. Sequence a second call only when
   its state or options depend on an earlier answer.
6. Route on uncertainty with per-action thresholds tuned on your own
   data. Ranking scores **order** (fail open); decision scores
   **authorize** (fail closed). Do not threshold a listwise or affinity
   number as if it were P(permit). The same judgment can authorize a
   reversible path and must not authorize an irreversible one.
7. Ship a decision-design card (below) and the smallest falsifying
   experiment. Name the eval path. A card without one is incomplete
   (`references/validation.md#eval--hill-climb`). Record family, model,
   rubric, candidate-source, and policy versions. Keep questions,
   criteria, and thresholds in one reviewable module; store raw
   judgments separately from derived actions.

## Mapping index

| Familiar method | Judgment shape | Detail |
|---|---|---|
| Mental models across domains (not SWE-only) | EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/Norman/snap-fit; **extractable-from-state boundary map** (self-contained vs needs outside knowledge). **Apply queued 1047 follow-ons** (`notes.md` §88): **Replica honesty** / **Empty ≠ approve** / **Beam as control** / **Cache is the exact envelope**. Apply 1144 (`notes.md` §89): **Typed if** / **Shadow then honor** / **Human every action** / **Atom then sense** / **File by Choice** / **Question preflight** / **Inbox read-only vs write**. Apply 1241 (`notes.md` §90): **Observe→score→act (namesake lock)** / **Decision-as-ranking** / **Native vs schema-guided Harbor** / **0 promotions / authored vs real** / **Offload + classifier-not-generator** / **Collapse late** / **Unofficial toolbelt** / **Pointer shell** / **Preview-first VOI / rubric rewrite** / **Life fail-open covers** / **S1 decide / S2 plan** / **Seed/expand/judge/verify + local daemon ≠ Jev**. Apply 1347 (`notes.md` §91): **Judge harness as control API** / **Batch packing VOI** / **Calibration as product** / **Decision-as-Plugin** / **Policy-constrained skill select** / **Tiny local econ pruner** / **Observe→score→act cousins** / **Deterministic verify ≠ System One** / **Soft-score vs hard-argmax**. Apply 1441 (`notes.md` §92): **Self-hosted econ** / **Soft-judgment gate integrity** / **Retrieval as calibrated decision space** / **Enterprise reflexes** / **Screenshot-free CU** / **Hybrid S1/S2** / **Harbor-jevals / SRE** / **Laya densifies** / **Demos / unofficial toolbelt**. Apply SIGNAL jevcache/jev-align (`notes.md` §93): **Decision ledger / memoization** / **GEPA alignment loop** (fingerprint after redact; recall vs decide; publish fingerprints+answers; CI replay as Harbor cousin; Cache hit ≠ correctness; hyperspaceai/jevcache ≠ kushals256/jevcache; human labels only; score never auto-accepts; production capture flywheel; sutro-sh/jev-align ≠ caiovicentino/jev-align). Apply SIGNAL enzyme/JA ModernBERT/Gemma (`notes.md` §94): **Compile-time System One / questions-as-index** / **Unofficial JA ModernBERT cross-encoder** / **NAR class legitimacy / multimodal observe→decide / Router-OOD** (guidance ≠ hook; catalysts ≠ summaries; compile-time System One; unofficial ≠ TypeSafe; format_version modernbert-jev/1; Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev; LFM default ≠ ModernBERT backend; Nemotron ≠ TypeSafe Jev; not a calibrated replacement; djev-dev complements djev-spark; images as Choice options; Laya essay numbers *theirs*; Router/OOD confidence; hosted bootstrap ≠ silent TypeSafe). Apply 1541 (`notes.md` §95): **Decision-as-plugin** (difficulty + policy thresholds + JSONL trace; jev-codex-pilot model + reasoning depth; keep/shadow/hybrid/reject) / **Evidence projection** (quarry evidence projection) / **Soft judgment integrity** (jevguard calibrator/cache/escape; jev-ci-selector CI shadow mode) / **Physical/control** (Frank-ZY-Dou/awesome-jev robotics/3D/control) / **Harbor-jevals injection-firewall** (one-dollar-tahoe TypeSafe Jev defense eval) / **llama.cpp replica** (llama-jev llama.cpp replica; petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator; seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard; webNeat/llama-jev ≠ WiktorB2004/llama-index-jev) Apply 1639 (`notes.md` §96): **OpenCode stdout-prune host port** (OpenCode jev-pruner context sieve; observe→score-candidates→prune; jev-zen / jev-1.13-free; zen-chat ≠ Noul; fail-open original; keepScore >0.1 floor; host port of tamaratran/jev-pruner; indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode; jev-webagent-bench empty stub; Kiln-AI/jev_jsonschema noul_threshold 0.5; NSStudent/JevSwiftSDK unofficial). Apply SIGNAL gliner-native-runtime (`notes.md` §97): **GLiNER2 native Apple path** (unofficial Swift/Core ML GLiNER 2.5-small; entity spans + confidence; not Choice/Score/Noul; not TypeSafe; label descriptions as schema; on-device ANE economics; honesty locks; shershah1024/gliner-native-runtime ≠ Fastino; ≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx; default threshold 0.1 still soft). Apply 1740 (`notes.md` §98): **Decision Graph Protocol envelope** (Decision Graph Protocol frame→assess→commit; app retains permissions/effects; Jev-first assessor-neutral; guarded commit / receipt/next frame; assessment batching; hard-gating DGP as safety theater; numerous-com/dgp ≠ TypeSafe official) / **Calibrated meaning-grep live tree** (jegrep calibrated path+range Nouls; no embeddings/index/daemon; ~$0.01–0.03 typical; agent --json; can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep) / **Archer-arch fidelity + measured calibration gap** (Archer-arch fidelity; kev family OOD 0.76–0.77 vs Jev 0.86; block-causal isolation; pointer/readout CE-trained; /v1/systemone drop-in; replica honesty). Apply 1843 (`notes.md` §99): **Cost-derived YES/NO/UNSURE control flow** (cost-sensitive decision theory × System One probabilities → control flow; thresholds derived from costs not hard-coded; YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human; auto-batching same-object questions; Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch) / **Typed-callback control flow** (judgment vs generation; deterministic execution after probabilistic judgment; exactly one app-owned callback; explicit uncertain branch; Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit) / **Variable-N option scoring as the trainable object** (dynamic candidate bags not fixed label sets; zwliJay/jev-forge ≠ NanoJev; not a new class-table species) / **Open NAR replica economics** (NAR local drop-in; open replica economics / latency vs closed Jev; wfzyx/von late-catch HIGH; competing NAR claims / replica honesty) / **Typed vs chat judges on guardrailing** (typed judgments vs chat judges on guardrailing; ishaannk/llm-vs-jev cross-note only; deeper integrity fold is rh-guard; nothing wins outright; can be argued out of guarding). Apply 1943 (`notes.md` §100): **Jev IS the if-statement** (judgments/probabilities drive branches; text model only writes prose; interpreter owns variables/loops/budgets/replay; otherwise maybe / confidence gate; chaos samples after the gate; southpolesteve/probably ≠ carldaws/hunch ≠ feelings ≠ Kungie/gut ≠ Illusion47586/judge ≠ tidymodels/probably) / **GEPA live delta** (133★ / forks 10 live; build calibrated classifiers from human feedback; HEAD/README SHA unchanged vs §93) / **Memory retrieve vs lease** (retrieve by relevance not resemblance; one calibrated yes/no per memory in one request; pointer mode 17/18 19/20 *theirs*; embedding resemblance misses the allergy; samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate; memory leases ended by new evidence; six Nouls then fixed rules in code; 0 of 157 false invalidations; questions/plans/directives are not evidence; unsure → review queue; host keeps the store) / **Contract-of-artifact lint rename** (name↔body / comment truth / test-claims; mizchi/jev-lint is mizchi/jevlint rename; no shipped rule has severity error; ~1 in 5 findings wrong *theirs*; mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint) / **JSON Schema question compiler** (JSON Schema → typed JSON via Jev; noul_threshold 0.5 decoder not a proof; IncompatibleSchemaError lists every bad property) / **Local System One economics** (on-device Laya CoreML ANE; ~5 ms P50 short decisions; 189/189 FP16 checkpoint parity; 10× not achieved; mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya; softmax over allowed tokens ≠ Noul; question-first cache; Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge; Jev-first Pi agent loop; slow-LLM fallback; explicit action menu / CandidateSource unimplemented; 62 tests wiring not quality; direwolfiy/JevPi ≠ standardagents/jevpilot ≠ pi-jev-control). Apply 2041 (`notes.md` §101): **Resume-screening bias audit** (resume-screening bias audit methodology; name×resume factorial independent Nouls; callback determined by resume quality; mean-probability name gaps operationally negligible; natemoo-re/bias-bench ≠ BBQ) / **MCDA panel code-owned verdict** (Plan/PRD panel → code-owned pass|review|block; cheerleading out of scope; austindixson/planalyzer ≠ single-goodness Noul) / **EU cost-aware routing** (cost-aware multi-model routing/escalation; decide vs do; successful-task cost; cannacre8ive/switchboard-ai ≠ ha-switchboard ≠ hermes-switchyard) / **Frozen-protocol bake-off** (frozen-protocol zero-shot bench; TypeSafe Jev vs PrismNLI vs Laya; contamination caveat; elcronos/jev-vs-open-decision-models ≠ JevBench ≠ DMB) / **VOI admission** (context-window admission control; VOI gate which tokens are worth the expensive model; fail polarity per lens; on small inputs lenses lose money; cvsgireesh/jevusher ≠ jev-sift ≠ winnow) / **Leveson control plane** (typed decision control plane; receipt ≠ authorization; historical-v0 zero retained cases; MokiMeow/jev-fabric ≠ jev-forge ≠ dgp) / **Scoring economics** (live 15-dim typed rubric re-score per pause; scoring economics exemplar; OpenJev/Codiv ≠ TypeSafe hosted; jose-troche/live-rubric ~$0.000004 desc / ~$0.000006 README) / **Pre-registered calibration science** (adversarial pre-registered Jev eval; 28 predictions before data; 123,805 requests; confidence does not track ignorance; polite injection 65% / crude 0%; willkelly/jev-evaluation ≠ jevals ≠ jev-baselines-eval) / **Class infrastructure SDK** (provider-neutral Elixir/BEAM Noul/Choice/Score SDK; class infrastructure; nshkrdotcom/system_one_sdk ≠ typesafe_sdk ≠ dannote/jev) | `references/mental-models.md` |
| Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, **domain specialist LoRA on independent gold**, openjev-lm, kev. Compaction job is backend-agnostic (Jev Score/Noul vs GLiNER2.5 encoder). Indexer cousin: GLiNER extract + escalate-S2 (10–50× unfilled). Computer-use observe→score-among-candidates→code-acts is backend-agnostic (Jev Ultrafast ↔ GLiNER2 Ultrafast ↔ Cua-S1 specialist ↔ Stagehand experimental Jev stack; **OCR+AX desktop:** typesafe-computer-use hosted Jev, never ships a screenshot for the *decision*; Cua-S1 is not TypeSafe Jev; Stagehand pick is a fast path, not a replacement; **≠** jev-macos-loop OmniParser **≠** camoufox; **ASR voice-browser:** jev-voice-browser hosted Jev, never ships a waveform). GLiFormer encoder serving `/v1/systemone` is a class-backend (jeff), not a Jev replica. Local MLX PCD is O(1) constrained-AR speed, **not** a calibrated Noul (system-one-benchmark Brier 0.3884 vs Jev 0.1096). **jevmlx** is the productized Apple Silicon one-pass schema→JSON+probs library (softmax ≠ Noul; no local leaderboard yet). **openvons** is an independent open-Jev class (LM/vision/voice finite-choice+prob; JevPick 3.2–4.8×; Flutter on-device; `/v1/systemone` wire-compat, not a TypeSafe replica). **OpenJev** (IamBusy) is a local 0.6B LoRA+scalar head on `/v1/decide` (45/60 *theirs*; **not** a TypeSafe drop-in; distinct from hraness/sysone OpenJev runners). **semif-serve** puts SemIf behind `/v1/systemone` (runoff, no option ceiling; 1164 vs 178 ms *theirs*; wire-compat ≠ replica). **grande** Rust/WebGPU kev-shaped branches (JGLUE *theirs* JNLI 0.614 ECE→0.088 / JCQA 0.853; 270M 0.710/0.710; isolation 0.098/0.996). **laya-jolt** Clojure/Jolt byte parity vs Python Laya. **JEV-CPU** SemIf on CPU (leesk212; Meanblock 404). **local-jev** ONNX ModernBERT measured not-equivalent (done 30% / shape 57% *theirs*). **GLiNER2→Choice/Score/Noul spec** (Eran-BA; design only, ≠ jeff GLiFormer). **openJev-verdict-2.0** competing NAR claims as **audit object not endorsement** (77.10%/0.0636/0.0144 *theirs*; PR #1; ≠ IamBusy/OpenJev `/v1/decide`). **chakuho** 1-token logprob local `/v1/systemone` (softmax ≠ Noul; coverage ≠ correctness; GUI 336 *theirs* 27B 95%/92% vs Jev 89%/82%). **jevinf** open replica engine (NanoJev/decider-2b/Laya; 2.57×/2.27× 100% argmax; MPS only). **laya-multilingual** mmBERT-base 322M (MASSIVE 0.366/0.387 vs English laya 0.227/0.733; Khmer 0.000@0.952 conf; ships uncalibrated). **schema-scorer** DeBERTa-v3-large scalar head (Hub; GitHub 404; v2 Choice 0.841 *theirs*; peaked ranking ≠ calibration). **githubnext/localjev** prompted-JSON `/v1/systemone` (MIT **261★**; TypeSafe SDK drop-in; wire-compat ≠ logit-equiv vs razorback16 structured-read; **≠** kunchenguid/local-jev; 1,200-req bake-off *theirs* Qwen3.6 76.7% / Gemma 26B 75.0% / DiffusionGemma 74.2% short; not calibrated). **NandhaKishorM/laya** PyPI+Router packaging of Hub Laya (Apache-2.0; **710★**; not a new species; T4 32.8 ms *theirs*; post-T ECE 0.081 vs Jev 0.246; Banking77 0.425 vs Jev 0.870; 0.766 is fine-tune not zero-shot; 0.85 still soft; **≠** TypeSafe `/v1/systemone`). **external openjev census** (@airesearch12 tweet ≠ jevbench v1.1; GLiNER2+routers class-boundary; incomplete vs watch). **JevBench v1.2 scored board** (geo-mean I/C/S/K 25% each; Jev 75.3 / SemIf 74.6 *theirs*; cal ON rank; instruction models in the table; ≠ v1.1 87.6; ≠ tweet census; Laya absent gap; Qwen3.8 27B ≠ Archer). **Hourly 0842:** already-folded class as a recipe (wire≠logit · product+FALLBACK · packaging honesty · pointer-not-generator · leaderboard VOI); skip thin noise. **Open LoRA replica, different jeff:** GestaltLabs/Jeff-1 LoRA Qwen3-4B ≠ logan-markewich/jeff GLiFormer; acc/ECE tradeoff n=9730 *theirs*; set reused. **Hourly 1241:** typesafeai-sdk-community not a new species; 2389-research/judgement license null; confidence ≠ winner p; jevbrain AUTO_ACT is not a Noul. **Hourly 1347:** FrancoisChastel/jev-code ≠ npm jev-code; ctmx/openrouter-jev-mcp Decision-as-Plugin; claudecode-jev-marketplace fail-open not hot path; pedroknigge/mcp_jev packs not ask_jev; 0thernet/system-one-skills deterministic verify; hari007sh/jev ≠ dannote/jev; nanoprune 2.8MB ECE 2.58%. **Hourly 1441:** Foq ~25ms/2.2GB local; rev prefill-only + HF jev-0.5b; robfrase/jev planning memo; meldltd/meldecision laya-go ONNX; laya-doom never pixels; logixism/laya-api empty README; akpsahan/laya ≠ Archer; Nibir1/typesafe-go ≠ official. **SIGNAL §94:** unofficial ≠ TypeSafe JA ModernBERT (argos1111/modernbert-ja-310m-jev; format_version modernbert-jev/1); Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev; LFM default ≠ ModernBERT backend; Nemotron ≠ TypeSafe Jev; djev-dev complements djev-spark (images as Choice options); Laya essay ≠ new species; hosted bootstrap ≠ silent TypeSafe. **Hourly 1541:** llama-jev llama.cpp replica; softmax ≠ Noul; **≠** TypeSafe **≠** pcdServer. **SIGNAL §97:** GLiNER2 native Apple path; unofficial Swift/Core ML GLiNER 2.5-small; entity spans + confidence; not Choice/Score/Noul; not TypeSafe; label descriptions as schema; on-device ANE economics; honesty locks; shershah1024/gliner-native-runtime ≠ Fastino; ≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx; default threshold 0.1 still soft. **Hourly 1740:** Decision Graph Protocol frame→assess→commit; app retains permissions/effects; Jev-first assessor-neutral; guarded commit / receipt/next frame; assessment batching; hard-gating DGP as safety theater; numerous-com/dgp ≠ TypeSafe official; jegrep calibrated path+range Nouls; no embeddings/index/daemon; ~$0.01–0.03 typical; agent --json; can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep; Archer-arch fidelity; kev family OOD 0.76–0.77 vs Jev 0.86; block-causal isolation; pointer/readout CE-trained; /v1/systemone drop-in; replica honesty. **Hourly 1843:** variable-N option scoring as the trainable object; dynamic candidate bags not fixed label sets; zwliJay/jev-forge ≠ NanoJev; not a new class-table species; NAR local drop-in; wfzyx/von late-catch HIGH; open replica economics / latency vs closed Jev; competing NAR claims / replica honesty; gut/judge are control-flow overlays not species. **Hourly 1943:** softmax over allowed tokens ≠ Noul; on-device Laya CoreML ANE; JSON Schema → typed JSON via Jev; noul_threshold 0.5 decoder not a proof; mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya; Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge; Jev IS the if-statement is a language primitive not a new species **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/judgment-class.md` |
| Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer; decision-token LoRA; packed one-forward logprob on open LLMs; **CUDA/PyTorch local replica** jevify — uncalibrated likelihoods ≠ Noul) / trained decision-only (Laya + ONNX port, Nimble, kev, **blackwood-rlcd** multimodal now, Archer Watch still Watch). **Domain LoRA specialist on independent gold** (not a Jev teacher-copy): train when downstream reads p; few-shot hosted when only argmax. Local `/v1/systemone` surfaces: jev-local (stub until `hf`), kev (trained pointer), von (§49 Needle SAN snapshot ≠ this-pass 395M / n=78; not a replica), **jeff** (GLiFormer-400M encoder, typesafe-sdk drop-in — not a Jev replica), **local-jev** (ModernBERT approximation — not equivalence). Laya ONNX: Mattepiu port vs **gqgs** complete browser int8 (distinct). Loopback **gateway** (sysone) routes hosted + local; does not run weights. Softmax over allowed tokens ≠ Noul. **jevmlx:** MLX one-pass schema→JSON+probs (Apple Silicon replica economics; not a Jev replica). **openvons:** independent open-Jev class (Apache-2.0 code; GitHub SPDX NOASSERTION); wire-compat `/v1/systemone`; NOTA + execute/confirm/reject. **OpenJev** `/v1/decide` ≠ TypeSafe. **semif-serve:** SemIf runoff wire (MIT pyproject / GitHub SPDX null). **grande** Rust/WebGPU `/v1/systemone` (license null; softmax ≠ Noul until T). **laya-jolt** Clojure Apache-2.0 byte-parity Laya. **JEV-CPU** CPU SemIf. **local-jev** ONNX NLI approximation (confidence omitted). **chakuho:** 1-token logprob constrained-AR endpoint (uncalibrated; coverage is format-mass). **jevinf:** Jev-kind engine + wire (not a replica). **laya-multilingual:** mmBERT-base for non-English; route by script. **githubnext/localjev:** Bun Chat Completions bridge; prompted JSON + entropy confidence; **≠** kunchenguid/local-jev; **≠** razorback16 logits. **NandhaKishorM/laya:** PyPI `laya` + Router over the three Hub ckpts; packaging ≠ new species; Jev still leads >20 options / soft-acc / raw ECE. **@airesearch12 census:** list ≠ rank; GLiNER2/routers counted as openjevs are a class-boundary. **JevBench v1.2:** same class table includes Luna/Gemini/DeepSeek/Qwen3.8; OpenJev on board = razorback16 DiffusionGemma ≠ IamBusy; SemIf formerly OpenJev. **≠** GestaltLabs/Jeff-1 Qwen3-4B LoRA replica. **SIGNAL §94:** unofficial JA ModernBERT cross-encoder (format_version modernbert-jev/1; unofficial ≠ TypeSafe; LFM default ≠ ModernBERT backend); Nemotron ≠ TypeSafe Jev / not a calibrated replacement; djev-dev complements djev-spark (native image / images as Choice options); Laya essay numbers *theirs* / Router/OOD confidence. **Hourly 1541:** llama-jev llama.cpp replica; softmax ≠ Noul; **≠** TypeSafe **≠** pcdServer. **Hourly 1740:** Archer-arch fidelity; kev family OOD 0.76–0.77 vs Jev 0.86; block-causal isolation; pointer/readout CE-trained; /v1/systemone drop-in; replica honesty. **Hourly 1843:** NAR local drop-in; wfzyx/von late-catch HIGH; open replica economics / latency vs closed Jev; competing NAR claims / replica honesty; variable-N option scoring as the trainable object; dynamic candidate bags not fixed label sets. candidate bags not fixed label sets; zwliJay/jev-forge ≠ NanoJev. **Hourly 1943:** on-device Laya CoreML ANE; ~5 ms P50 short decisions; 189/189 FP16 checkpoint parity; 10× not achieved; softmax over allowed tokens ≠ Noul; question-first cache **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/judgment-class.md` (when-to-use table) |
| Entropy as allocator (low / medium / high) | Typed low+medium decisions → System One marginals; high-entropy synthesis → frontier decoder. Product rhetoric, not a meter. **Hypothesis** | `references/judgment-class.md` |
| Formal / semi-formal (proof vs judgment) | Sensor vs constraint vs searchlight; Alloy vs Apalache; DST trio; TOCTOU-of-Noul, AI×FM. Skills→oxlint: AST/precheck compose with remainder judgment without hard-gating a Noul as a proof. PR attention ≠ correctness (anti-soundness-theater). Engine owns truth / Jev owns judgment (Stockfish+Jev chess coach). Capability kernel: type-safe ≠ correct; irreversible behind threshold AND human. Eval integrity: check the instrument, not just the score (`dinostomp jev` tests a question like an if-statement). Effect contracts, not surface tokens (construct-auto-classifier; privilege ≠ verdict). **Jev supplies evidence, code owns authority** (actiongate-jev; a positive score never overrides a deterministic security failure). **Turnstile clone:** deterministic policy + Jev remainder + replay (evidence ≠ authority). **Type-safe ≠ correct as jaggedness receipts** (atlas; schema-valid ≠ picked-right). **TLA+ compose with a Jev-class oracle:** never confidently wrong; escalate is the safety valve (jev-labs; inverse of soundness theater is hard-gating without escalate). **Hourly 1347:** typed-gate band [0.40,0.60] is refusal; pi-jev-gate fail-closed; choice is the verdict; jev-calibration-arena never acts. **Hourly 1441:** typesafe_agent_gates 27/27 / 31/31; EpicEric/safe-sh static remainder; pastepilot Confirm before act; typesafe-scheduler-diagnostics advisory; choxos/jevchess engine owns truth. **Advance/coverage ledger:** Jev answers questions; SEAL answers whether the world may change (coverage.path auto|code|human|escalate; mint ≠ product brain). **Conflict ≠ ignorance:** Noul collapses both; named Choice escape separates (typed-evaluation-collapse; schema-as-interface). **Sentence-as-rule lint:** ast-grep matcher silent × Jev `ask:` loud (mizchi/jev-lint is mizchi/jevlint rename; ≠ huntedman/JevLint). **Whole-repo intent:** VERIFIED/VIOLATION/UNKNOWN; empty search ≠ proof (jev-intent-review). **Decision-as-assert:** meaning Noul vs exact `toContain`; ambiguous band fails both polarities (jevtest; 0.85 still soft; record/replay; hard-gating a matcher as a merge seal is soundness theater). **Authorship named escape:** `human`/`ai_generated`/`uncertain`; not courtroom evidence. **Empty findings as approval, or auto-promoting an agent-written workflow, is the same theater** (stanley-code `notChecked`; Soft Noul ≠ hard safety). **feelings `.feels()`** default 0.5 never rounded; **Essentiel-Jev never authority**; **enzo-mcp UNKNOWN**; **pigeonhole OTHER skip**. **Hourly 1241:** ZHUBoer/ego-jev reserved `__none__`; runWorkflow completed ≠ success; ORIGIN pause-if-no-Jev; jevbrain AUTO_ACT is not a Noul. **SIGNAL §93:** Cache hit ≠ correctness; score never auto-accepts. **SIGNAL §94:** guidance ≠ hook; unofficial ≠ TypeSafe; hosted bootstrap ≠ silent TypeSafe; Nemotron ≠ TypeSafe Jev; not a calibrated replacement; Router/OOD confidence. **Hourly 1541:** jevguard calibrator/cache/escape; jev-ci-selector CI shadow mode; one-dollar-tahoe TypeSafe Jev defense eval (rh-guard owns the gate cousin). **Hourly 1740:** Decision Graph Protocol frame→assess→commit; app retains permissions/effects; Jev-first assessor-neutral; guarded commit / receipt/next frame; assessment batching; hard-gating DGP as safety theater; numerous-com/dgp ≠ TypeSafe official. **Hourly 1843:** thresholds derived from costs not hard-coded; cost-sensitive decision theory × System One probabilities → control flow; hard-gating 0.038 as safety theater; "guaranteeing" calibration is theater; deeper integrity fold is rh-guard. **Hourly 1943:** noul_threshold 0.5 decoder not a proof; no shipped rule has severity error; 0 of 157 false invalidations; questions/plans/directives are not evidence; IncompatibleSchemaError lists every bad property **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/formal-methods.md` (one-screen: `references/formal-semi-formal.md`) |
| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured); **Harbor-shaped cousin:** decide→policy→LLM leftover (shared Answer schema; jev vs gen-json vs gen-logprob; Noul 0.5 never rounded). **Closed-vote CU:** no planner LLM; code builds options, Jev only picks (JevOnly). **Host-owned product:** app retains handlers/permissions; Jev over live typed actions (waymode). S1 specialists + S2 coordinator is the same split (description-only greenfield). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision. Fail polarity is per act: skip-wake fail-open vs merge-gate BLOCK fail-closed; OMP/pi acceptance+route **fail-open** (contrast pi-jev-approver fail-closed). OMP prompt suppression: operator owns the bar; not a sandbox (omp-greenlight). Skill-broker outline: Jev never grants access. Hybrid local decide + remote fill; `DONE` ≠ verified success. Specialist computer-use: plan ≠ execute, dry-run default (Cua-S1; not TypeSafe Jev). Judgment as a language primitive (Ruby `almost_certain?`/`pick`/`rate`). Decision-native RAG: retrieve wide → decide → evidence set → LLM. Classify-first MCP (topology A): content to the judge without entering main agent context first. Generative UI: model decides, compiler emits. Draft-gate silence ≠ safer (heartbeat). Living class-pattern atlas (not a 342-title dump). Stagehand experimental Jev: pick-and-copy extract + act tree + observe/cache-check; LLM fallback; draft stack. Public judgment wall (six parallel questions; policy-in-code; cost-to-1M). PR attention ≠ correctness. Session-sticky first-prompt route (fail-closed fallback). Capability kernel (LLM ring 3 / Interlock ring 0; secrets never in agent; Jev SENSOR; policy.py BLOCK/ASK/ALLOW; type-safe ≠ correct). Typed control plane around DSPy (drafts AFTER route+action). Engine owns truth / Jev owns judgment. Human-confirmed kill (mapped explanations; identity re-check). Wire-compat encoder backend (GLiFormer `/v1/systemone` drop-in; cheaper, less accurate on reasoning-heavy). Loopback gateway routes hosted + local (not a model). **Constrained optimizer + S1 features** (slo-router: Jev never the sole hot-path gate; fail-open local features). **Effect-based shell gate** (construct: privilege ≠ verdict; fail-closed). **Attention filter / VOI** (jev-lens: never blocks the agent; never green unless sure). **Jev supplies evidence, code owns authority** (actiongate-jev). **Measurement owns endorsement** (jev-packs evidence-gated). **Ranking ≠ calibration** (does-jev-confidence; never hard-threshold raw p). **Hot-click CU** (ego-jev: indexed table → operation+target; text model only for type). **Jev judges relevance, code decides structure** (jev-compactor; never rewrite; regex floor). **Local rules first / never auto-train on own hides** (x-reply-filter). **Control-plane combinators** (Then/Gate/Vote/Cascade/Weighted; not chat turns). **Skill VOI / abstention** (skillranker; hook fail-open). **Receipts not leaderboard** (atlas + frontier-100 + OOD). **Turnstile** evidence≠authority + replay. **MLX one-pass replica economics** (jevmlx; softmax ≠ Noul). **TLA+ consensus kernel** (jev-labs; never confidently wrong; escalate). **SEAL advance/coverage** (no seal, no advance; exception queue visible). **Sureness bands** (how-sure-is-jev; max_prob is generous). **JevBench v1.1** (calibration reported, not scored). **CI typed gate** (ci-gatekeeper before expensive review). **Codex MCP adapter** (jev-in-codex; ranking unbenchmarked; lexical fallback). **Stop-hook attention redirect** (jev-preflight; eight axes; assist=one reinspect; fail-open; not a merge blocker). **Pre-send view selection** (dizk/jev-lens; 79% fewer tokens *theirs*; compress-before-first-send; distinct from rashedInt32/jev-lens). **Observational memory** (pi-om; keep/kind verbatim; model-free compact). **tools≠use** (carryforward 0/4 recall; SessionStart > hoping). **Physical-world S1** (HA-Jev; sensors from typed answers; not for locks/heaters). **Judgment outside the store** (jevql CLI; DB never sees `jev()`). **Landed-script / headless≠auto-approve** (construct); **digital-design combinators** (jev-combinators rename + Router/Loop/Retry/Fallback/Memory; metaphor ≠ literal AND/OR); **VOI cache admission** (jevcache same-intent; 0 FP/100 *theirs*; fail-open); **worth-your-attention VOI** (ThinkyMiner/Winnow 80%/90%; ≠ kevinpita/winnow); **Jev WHETHER / Python HOW / LLM WHAT** (hermes-jev-router; license null); **typed escalate/continue/abort baton** (jev-handoff; inverted loop; gate never grants; fail-open); **Playwright executes, Jev chooses** (browser-jev; sample-from-distribution); **OpenJev `/v1/decide` ≠ drop-in** + **SemIf runoff wire**; **conflict ≠ ignorance** (named Choice escape); **decision-as-memory flywheel** (DGUI_HYPERMEM-JEV 6-row schema); **record/replay CI** (jevassert landed; accuracy+ECE+cost gates offline); **failure-finding arena** (chenmingtang830/jevarena ≠ meetr1912/jev-arena); **BBQ** 97.28%/0.04/0.34/$0.3429 *theirs*; **decider≠executor** (jeffrey: Jev next-tool, LLM fills args); **sentence-as-rule** (mizchi/jev-lint is jevlint rename); **VOI hunk prune** (prune-review ~20% target; 1.18% with outlier); **persist constraints across compaction** (pi-heed); **Harbor SGR-judge contract** (jev-judge-bench; canaries ≠ quality; ≠ jevarena/jevbench); **hand no-text steps** (jev-use; Vercel drops confidence; margin 0.4; fail-open gate; ≠ jev-ultrafast); **Pi System-One control plane** (pi-jev-control; GUI never force-click; compaction never writes session); **never free-generates** (jev-gpt tree of Choices; 400 calls / 75 s / 2¢ *theirs*); **OpenRouter recipe atlas** (jev-cookbook; 16–36 item samples not benches; 425 calls / $0.015); **personal-history feed** (jevfeed; no social graph; one request per batch of ten); **competing NAR claim-audit** (openJev-verdict-2.0; dual-channel ECE; PR #1; ≠ IamBusy/OpenJev); **empty compaction-proxy skip** (IPECTER context-pruner **and** jev-runway LICENSE-only); **1-token logprob endpoint ≠ Noul** (chakuho; coverage ≠ correctness); **open replica engine** (jevinf argmax-parity); **unofficial Elixir SDK ≠ OTP peer** (typesafe-elixir-sdk ≠ dannote/jev); **jevex n=16 files-to-read VOI** (rename of jev-semantic-explorer); **commit pre-review attention≠verdict** (commitjev; middle band never rounded); **Hermes plugin is Agnes not TypeSafe** (hermes-plugin-jev); **pi-jev-compact ≠ pi-jev-compaction** (verbatim summarizer replacement); **decision-native inbox** (mailordinal; humans own ambiguity); **unofficial jev-cli not ready** (≠ jevql); **laya-multilingual** English checkpoint confident-wrong OOD; **schema-scorer peaked ranking ≠ calibration**; **productized System One HTTP** (classifier.dev; label+p; batch ~1000; Jev primary / LLM fallback); **escalate-under-threshold** (smart single-label <0.7; multi-label ignores); **silent FALLBACK** (granite 0.546 vs advertised 0.800; rh-guard owns the gate); **systematic-review pointer (choxos/jev-reviewer ≠ egma-ai)** two-pass Choice+Noul; *Not found* is an answer; human check is the product; **githubnext/localjev** prompted JSON ≠ structured-read logits (wire-compat ≠ logit-equiv; **≠** kunchenguid/local-jev; GitHub Next **261★**; 1,200-req caveats *theirs*); **NandhaKishorM/laya packaging** Router script-before-p; post-T ECE ≠ raw ECE; 0.85 still soft; Banking77 token-budget; **≠** TypeSafe drop-in; **external census ≠ scored bake-off** (@airesearch12; GLiNER2+routers class-boundary; incomplete vs Laya/localjev/kev; Harbor honesty watch); **JevBench v1.2 geometric-mean product** (I/C/S/K 25% each; cal ON rank; Luna I=96.8 rank #7; ×2/est. Harbor honesty; option-order 72→21; Laya absent gap; Qwen3.8 27B ≠ Archer); **hourly 0842 apply-the-five** (already §73–§78; do not re-card); **skip thin noise**; **hard-gate a Noul as a PR/quality gate is soundness theater** (totally-tim/jev-gate / claude-jev-warden; ≠ jev-gateway / MongLong0214/jev-gate / jev-gate-student-b); **S1 keeps flying / S2 one-use** (khordoo delta: escalate without stall; purple = consumed; Local controller ≠ githubnext/localjev; seed = geometry; 20% still soft; no pixels; S2 never grants). **OCR+AX desktop CU:** typesafe-computer-use (hosted Jev; never screenshot-to-frontier for the decision; overlapping options = doubt; writer/decider; 155× *theirs* one screenshot; 0.4/0.5 still soft; **≠** jev-ultrafast **≠** cua-s1 **≠** camoufox). **ASR voice-browser CU:** jev-voice-browser (partial-speech VOI; pointer spans; spoken confirm ≠ auth; 27/27 *theirs* fixtures; **≠** jev-voice-control **≠** nikolas-j **≠** OCR desktop). **Wrap-as-execution ALLOW/ASK/DENY:** wrap *is* the tool function; rules first; ASK throws; fail-closed (AgentGhost; **≠** actiongate **≠** jev-use fail-open; rh-guard owns the gate cousin). **JP genre atlas:** apps by hole; stars research-time; not verified evals (@studio_yebisu; **≠** class census §77 **≠** v1.2). **External pedagogy:** Akshay “Jev Clearly Explained”; LLM hammer; schema-safe ≠ correct; 200×/400× TypeSafe ceiling; shadow + questions-as-code; **≠** official docs **≠** Flavio **≠** AgentGhost. **Meaning-grep dedicated:** proposition≠embedding; boolean composition of thresholded Nouls; Semgrep.dev collision; not a gate (jev-semgrep §86). **Hourly 1047 + deferred 0945:** decision-validated UI (gram-render never authors text; jev2ui Jev decides / Gemini writes); decision-as-assert (jevtest ambiguous band 0.15–0.85 fails both; 0.85 still soft); hybrid S1 (anima3 closed verb menu + hard safety first; jeff confidently flat on magnitude; Qwen logprob default — do not invent Laya as a backend); pointer search (JevFind path then window); Harbor bake-offs three shapes (frontier-bench ≠ frontier-100; GLiClass product bakeoff not architecture duel; four engines / majority floor / calibration ≠ discrimination); authorship named escape (not evidence); non-SWE (ha-switchboard HA remains execution ≠ HA-Jev; n8n Low Confidence abstention); compaction delta (fast-jev-compaction-pi ≠ pi-jev-compact ≠ pi-jev-compaction); full-distribution optimizer (jevloop UCB1+CEM; no LLM in the loop; mock default); deferred class (laya-vision SmolVLM `score` untrained ≠ blackwood ≠ Archer; Cerebellum-2B `/v1/decide` ≠ TypeSafe — wire-compat vs agent-routing as separate Harbor axes, competing NAR not endorsement; laya-grounded not drop-in / phishing 0.611→0.512 / Platt not temperature). Hard-gating a Noul as test/PR/HA write/authorship seal is soundness theater. **Queued SIGNALs:** Collapse GestaltLabs/Jeff-1 into logan-markewich/jeff; Quote Jeff-1 ECE as “better than Jev”; Treat empty stanley findings as approval; Treat findme beam score as file-identity; Price workers, not the conversation. Soft Noul ≠ hard safety. **Hourly 1144:** Treat `.feels()` 0.5 as a bool if; Collapse apa-agent-harness into AntonioCoppe/jev-harness; Quote apa "mathematically fulfilled" / npm @aipersona; Treat grok-bot-jev A/B as token savings; Let Essentiel Jev send / skip human approve; Collapse enzo-mcp into jev-sift / skip UNKNOWN; Treat pigeonhole OTHER as a move / 0.6 as Harbor τ; Treat the HF playground as live Jev / collapse into classifier.dev; Quote jev-reliability as accuracy; Collapse clduab11/jev-test into realZachi/jevtest / paste bars as results; Paste "Jev wins" from jev-rag-benchmark; Collapse dairui1/jev-lab into BrendanH18/jev-lab / re-card jev-desktop; Treat jevmail as mailordinal / mailjay as read-only. **Hourly 1241:** Collapse ZHUBoer/ego-jev into jiangkoumo / treat `completed` as success; Treat jsort logits as frequencies / Choice as the scale; Paste groundedness Macro-F1 as “Jev wins quality”; Quote jev_playground 83% / promote from authored bars; Copy `jev-latest` on Zen; Treat techstack ranks as a generated stack; Collapse s1_ruby into hunch/feelings; Treat judgement as jevql / confidence as winner p; Treat the Rust community SDK as official / a new species; Collapse tpellet/hunch into carldaws/hunch / skip exit 3; Quote file-search 15 matches as recall; Treat linkmap referee as gold / let Jev see S2 prose; Treat jev-mail as jevmail / tidy OTHER as a move; Close pinned/audio/current tabs / skip Show; Let ORIGIN LLM decide / continue without Jev; Gate crawlers on raw `bug_likely`; Sell jevbrain AUTO_ACT as a Noul. **Hourly 1347:** hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev; cyrusasco/typesafe-mcp noul deadband 0.35–0.65; smartdio/jev-browser-agent ≠ ZHUBoer/ego-jev; Dakai/omp-jev-web DONE ≠ proof; typed-gate band [0.40,0.60] is refusal; pi-jev-gate fail-closed; choice is the verdict; Jev-Calibration Platt ECE 0.117→0.052. **Hourly 1441:** Jev-Reranker live Jev not yet measured; sessionwise opt-in relevance; jev-search pointer sieve; savka777/jev-search ≠ kazuhideoki/jev-search ≠ superagents-lab/jev-search; 400ms Salesforce WebMCP; droidjev screenshot-free; Tewoto1 jevcu planner still writes; ha-conversation-jev Jev→Grok; dsh-jev can only gate; jev-classification-benchmark specified not run; jev-luna-pagerduty p≥0.50; jev-drive sim not AV; story-arc Jev never authors; jev-hs-assistant HS6; golergka/jev-plays-starcraft-2 UI-verified ≠ API Victory; awesome-jev-use-cases catalog. **SIGNAL §93:** fingerprint after redact; recall vs decide; publish fingerprints+answers; CI replay as Harbor cousin; Cache hit ≠ correctness; hyperspaceai/jevcache ≠ kushals256/jevcache; human labels only; score never auto-accepts; production capture flywheel; sutro-sh/jev-align ≠ caiovicentino/jev-align. **SIGNAL §94:** guidance ≠ hook; catalysts ≠ summaries; compile-time System One; unofficial ≠ TypeSafe; format_version modernbert-jev/1; Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev; LFM default ≠ ModernBERT backend; Nemotron ≠ TypeSafe Jev; not a calibrated replacement; djev-dev complements djev-spark; images as Choice options; Laya essay numbers *theirs*; Router/OOD confidence; hosted bootstrap ≠ silent TypeSafe. **Hourly 1541:** difficulty + policy thresholds + JSONL trace; jev-codex-pilot model + reasoning depth; keep/shadow/hybrid/reject; quarry evidence projection; Frank-ZY-Dou/awesome-jev robotics/3D/control; one-dollar-tahoe TypeSafe Jev defense eval; jevguard calibrator/cache/escape; jev-ci-selector CI shadow mode; llama-jev llama.cpp replica; petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator; seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard; webNeat/llama-jev ≠ WiktorB2004/llama-index-jev. **Hourly 1639:** OpenCode jev-pruner context sieve; observe→score-candidates→prune; jev-zen / jev-1.13-free; zen-chat ≠ Noul; fail-open original; keepScore >0.1 floor; host port of tamaratran/jev-pruner; indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode; jev-webagent-bench empty stub; Kiln-AI/jev_jsonschema noul_threshold 0.5; NSStudent/JevSwiftSDK unofficial. **SIGNAL §97:** GLiNER2 native Apple path; unofficial Swift/Core ML GLiNER 2.5-small; entity spans + confidence; not Choice/Score/Noul; not TypeSafe; label descriptions as schema; on-device ANE economics; honesty locks; shershah1024/gliner-native-runtime ≠ Fastino; ≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx; default threshold 0.1 still soft. **Hourly 1740:** Decision Graph Protocol frame→assess→commit; app retains permissions/effects; Jev-first assessor-neutral; guarded commit / receipt/next frame; assessment batching; hard-gating DGP as safety theater; numerous-com/dgp ≠ TypeSafe official; jegrep calibrated path+range Nouls; no embeddings/index/daemon; ~$0.01–0.03 typical; agent --json; can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep; Archer-arch fidelity; kev family OOD 0.76–0.77 vs Jev 0.86; block-causal isolation; pointer/readout CE-trained; /v1/systemone drop-in; replica honesty. **Hourly 1843:** judgment vs generation; deterministic execution after probabilistic judgment; exactly one app-owned callback; explicit uncertain branch; YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human; auto-batching same-object questions; Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch; Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit. **Hourly 1943:** Jev IS the if-statement; judgments/probabilities drive branches; text model only writes prose; interpreter owns variables/loops/budgets/replay; otherwise maybe / confidence gate; chaos samples after the gate; Jev-first Pi agent loop; slow-LLM fallback **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/mixed-architecture.md` |
| Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key. Encoder cousin: GLiNER2.5 retention Choice + char-offset spans (gliner25-compaction); fail-closed keep_full; shadowMode default. Stdout cousin: jev-pruner (Jev Noul after hard ≤10k/JSON-diff envelope; fail-safe original; archive). Session-ledger cousin: carryforward (verbatim facts; Jev scores recall; rules never judged; fail-open dump). Classify-first MCP cousin: jev-sift (batch path/url/text → Jev without entering main agent context; uncertain/errors/truncation ≠ irrelevant). **Framework-agnostic compact+gate:** Jev judges relevance, code decides structure; never rewrite; regex floor in code; compaction fail-open if Jev down, safety gate fail-closed (jev-compactor later bench **73%** / 350 ms / 4 of 4 *theirs*; OpenCode port fast-jev-opencode already §62). **Pre-send views:** dizk/jev-lens (79% fewer tokens on 500 SWE-rebench trajectories; code full unless confident). **Observational memory:** pi-om keep/kind verbatim. **tools≠use:** carryforward SessionStart hook > MCP sitting there. **Hermes WHETHER/HOW/WHAT:** compact original chunks; skip next main-model when evidence is enough (hermes-jev-router; needs core patch; fail-open). **Empty compaction-proxy skip:** IPECTER/jev-context-pruner **and** IPECTER/jev-runway LICENSE-only. **Pi verbatim summarizer replacement:** pi-jev-compact (keep/drop tool calls; fail-open to LLM summary; ≠ vava-nessa/pi-jev-compaction). **Dedicated Pi port of fast-jev-compaction:** zaycruz/fast-jev-compaction-pi (verbatim keep/drop; fail-open to built-in LLM summary; ~50× *theirs*; **≠** pi-jev-compact **≠** pi-jev-compaction). **Atom then sense MCP** (enzo-mcp independently falsifiable claims; enzo-mcp UNKNOWN useful; **≠** jev-sift). **Hourly 1347:** 0thernet/system-one-skills deterministic verify; nanoprune 2.8MB ECE 2.58%. **OpenCode stdout-prune host:** indiejoseph/opencode-jev-pruner (jev-zen / jev-1.13-free; zen-chat ≠ Noul; fail-open original; keepScore >0.1 floor; ≠ tamaratran ≠ fast-jev-opencode; `notes.md` §96). **Hourly 1740 DGP ≠ this job** (protocol envelope around assessor; commit is the guard; `notes.md` §98). **Hourly 1843 gut auto-batch** is same-object fan-out, not a sieve. **Hourly 1943:** retrieve by relevance not resemblance; one calibrated yes/no per memory in one request; pointer mode 17/18 19/20 *theirs*; embedding resemblance misses the allergy. **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/applied-mappings.md#1-context-sieve` |
| Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt; char-offset compaction same species). Observed a11y/DOM controls: score among them; code clicks (Jev or GLiNER2 or Cua-S1 option-attention). Harness productization: Stagehand extract pick-and-copy (Jev picks; code copies; schema/gate else LLM). **Closed-vote CU:** code builds options, Jev only picks, no planner LLM (JevOnly); host-owned handlers/permissions (waymode). **Hot-click CU:** indexed element table → operation+target in one request; code owns observe/execute/verify; text model only for type (ego-jev; ~2× vs per-step LLM, n=3, not a bench). **OCR+AX desktop CU:** numbered OCR+AX items; hosted Jev; writer only for free text (typesafe-computer-use; exclusive actions; split kind/item/site; **≠** jev-ultrafast). **ASR voice-browser CU:** regex spans; Jev picks; code copies (jev-voice-browser; numbered overlay, no second model). Meaning-as-spec: Cucumber .feature only; Jev picks among observed controls (jevcumber). **Evidence-synthesis pointer (choxos/jev-reviewer, ≠ egma-ai):** two-pass Choice (which line) + Noul (does this line itself answer); *Not found* is an answer; human check never overwritten. **Adversarial browser:** Playwright executes, Jev chooses next act; sample from the distribution not argmax; fail only high conf **and** high severity (browser-jev). **Decision-validated UI:** derive candidates, Jev selects, compiler emits; Jev never authors text (gram-render Telegram GramSpec; jev2ui A2UI + leftover Gemini). **Pointer search:** path Noul then window; Jev picks files/line ranges; code copies (JevFind; 0.25/0.55 still soft). **NL memory → beam-search FS:** findme; **≠** JevFind path-then-window. **Decision-as-filing:** pigeonhole OTHER skip. **Observe→score→act namesake:** ZHUBoer/ego-jev reserved `__none__`; runWorkflow completed ≠ success. **Downloads fail-open:** tidy none-of-folders stay. **Hourly 1541:** quarry evidence projection. **Hourly 1639:** OpenCode jev-pruner context sieve (verbatim chunks + markers). **SIGNAL §97 ≠ this job** (schema→spans locate, not keep/drop of held candidates). **Hourly 1740 jegrep** returns path+range pointers (ranking/gather, not keep/drop of held candidates). **Hourly 1843 jev-forge** scores *supplied* candidate bags (selector, not keep/drop of held bytes). **Hourly 1943:** name↔body / comment truth / test-claims; mizchi/jev-lint is mizchi/jevlint rename; no shipped rule has severity error. **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/applied-mappings.md#2-exact-text-keep--drop` |
| Environment / harness triage | Scan every step for env failure; LLM autopsy only on flags. Merge-gate cousin: cluster in code, judge labels cause, policy owns PASS/BLOCK (latch; judge never says ignore alone). **Pre-review typed gate:** should_review/risk/route/touches_secrets → auto-approve|human-review|block (ci-gatekeeper; cheap before expensive LLM/human; operator-owned thresholds). **Stop-hook attention redirect:** eight risk axes, assist=one reinspect, fail-open, uncalibrated 0.85 (jev-preflight; not a merge blocker). **VOI hunk prune:** Jev scores PR hunks before expensive generative review (prune-review; 22-run 1.18% with 305% outlier *theirs*; ~20% target). **Whole-repo intent:** VERIFIED/VIOLATION/UNKNOWN beyond the diff (jev-intent-review). **Commit pre-review:** message vs diff / cohesion / omitted change; middle band is review not a verdict; regex proves literals (commitjev). **Hourly 1241:** jev-crawlers risk bands never raw boolean. **Hourly 1347:** alsoleg89/decide packing VOI; 0.8 ≠ 80% accuracy; 0thernet/system-one-skills deterministic verify | `references/applied-mappings.md#3-environment--harness-triage` |
| Moderation and ranking | Hold-before-publish vs graded rerank; fail policy per action. Meaning-search without embeddings (jevgrep packed parallel; 79% top-5 on stripped repos; keyword still wins exact strings). **Meaning-grep** AND/OR/NOT over *thresholded* line Nouls; proposition≠embedding; Semgrep.dev collision; not a gate (jev-semgrep §86). **Evidence-packet explorer:** index-once ask-many, citable source_of_truth/tests/callers (jevex). Measured RAG rerank vs generative rerank (Jev-RAG one-run ≥70% cost / 72% latency vs Spark rerank; full-context Spark still faster). **Local rules first then remainder Nouls:** never auto-train on the model's own hides (x-reply-filter). Minimal consumer labels: bohutang/sift ~$0.00003/post. **Worth-your-attention VOI:** ThinkyMiner/Winnow read/skim/save/skip from typed answers (80%/90% *theirs*; always qualify vs kevinpita/winnow). **Personal-history ranking without a social graph:** jevfeed (one Jev request per batch of ten; distribution *is* ranking; history never uploaded). **OpenRouter recipe atlas:** jev-cookbook (triage/PII/rerank/browser; samples not benches). **Decision-native inbox:** mailordinal (nine typed signals → 100-point policy; humans own ambiguity). **Evidence-packet explorer delta:** jevex n=16 SWE 160s→69s / $8.74→$3.13 *theirs* (rename of jev-semantic-explorer). **Productized classification API:** classifier.dev (label+calibrated p; spam/inbox/feedback; **185★**). **Open NAR packaging:** NandhaKishorM/laya PyPI+Router (**710★**; Hub weights; not HTTP Jev). **Pointer search:** JevFind path then window (0.25/0.55 still soft). **Authorship named escape:** `human`/`ai_generated`/`uncertain` (not evidence). **n8n classify/route/score:** Route by Choice + Low Confidence (unofficial; 0.5 still soft). **Four-engine Harbor:** job-posting-triage majority floor 0.947; fitted tfidf wins; calibration ≠ discrimination. **Read-only Gmail trays:** jevmail gmail.readonly; **macOS inbox writes after review:** mailjay. **Hourly 1241:** jsort scores are relative; Noul not Choice for scale; muhammedilyasy/jev-mail metadata only; lkclean Show fail-open; jev-yt-time-saver Show anyway. **Hourly 1347:** nanoprune 2.8MB ECE 2.58%; alsoleg89/decide packing VOI. **Hourly 1740:** jegrep calibrated path+range Nouls; no embeddings/index/daemon; ~$0.01–0.03 typical; agent --json; can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep. OpenRouter/TypeSafe auto-failover is silent FALLBACK. **Hourly 1843:** typed judgments vs chat judges on guardrailing; ishaannk/llm-vs-jev cross-note only; deeper integrity fold is rh-guard; nothing wins outright. **Hourly 1943:** ~1 in 5 findings wrong *theirs*; mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/applied-mappings.md#4-moderation-and-ranking` |
| Skill / tool routing | Choice over a closed catalog + whether-anything-fits; code dispatches. Route ≠ memory: cheap intent gate skips memory tours on easy routes. Session-sticky first-prompt classification (lock for the session; fail-closed to a declared fallback). OMP/pi: `jev_route` topology/tier + `jev_acceptance_gate` before done (**fail-open**; contrast pi-jev-approver fail-closed). OMP prompt suppression: Jev grades gated calls; operator owns the bar; plugin never self-tunes (omp-greenlight; not a sandbox). **Outline only:** Hermes pre-agent skill broker — code owns grants; Jev never grants access (skill-broker; not a production recipe). **Constrained optimizer:** Jev supplies task/exactness/evidence features; controller owns SLO/quality floors; fail-open local features (slo-router; measured p95 77.93→490.38 same routes). **VOI over skill library:** two-pass + none-of-these; hook never blocks (skillranker 52★). **Sibling contrast:** skill-broker grants in code (fail-closed foundation-only) vs skillranker advisory vs turnstile runtime authorize. **Codex MCP adapter:** jev_select_capability / jev_search / jev_triage; caller supplies catalog; lexical fallback (jev-in-codex). **Harbor roster-size harness:** BM25 vs Jev at 50–500 (pi-jev-skill-bench; 43 gold; no live numbers this pass). **Pi strip-roster:** two-stage gate 0.30 / fits 0.40; no key → no-op; tool mode is tools≠use cousin (pi-jev-skill-suggestion). **Pi System-One control plane:** pi-jev-control (task/model/skill/memory/review/GUI; no live quality numbers; license null). **Host-adapter surface delta:** jev-routing adds Cursor Agent CLI / Devin CLI (still not MCP). **Hermes plugin ≠ TypeSafe:** hermes-plugin-jev is Agnes 3.0 Flash chat-completions branded as Jev. **Competing NAR agent-routing:** Cerebellum-2B pointer over candidates on `/v1/decide` (**not** TypeSafe `/v1/systemone`; wire-compat vs agent-routing as separate Harbor axes; do not endorse vs-Jev 94.92%/81.1%). **Jev-first bounded agent:** stanley-code; Empty findings ≠ approval. **Price workers, not the conversation:** jevsubrouter; Stats are counts, not dollars. **Cheap decision layer / skill honor:** grok-bot-jev; **APA harness cousin:** apa-agent-harness **≠** AntonioCoppe/jev-harness. **Hourly 1241:** yuyang2230/jev-agent-skill jev-1.13-free; jev-techstack-classifier stack_config.json; tpellet/hunch exit 3. **Hourly 1347:** hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev; ctmx/openrouter-jev-mcp Decision-as-Plugin; claudecode-jev-marketplace fail-open not hot path; FrancoisChastel/jev-code ≠ npm jev-code. **Hourly 1541:** difficulty + policy thresholds + JSONL trace; jev-codex-pilot model + reasoning depth | `references/applied-mappings.md#5-skill--tool-routing` |
| Expensive observation router | Structural prove (text layer) ∩ remainder Noul (needs OCR?) | `references/applied-mappings.md#6-expensive-observation-router` |
| Capability kernel / human-confirmed gate | Secrets never in the agent; closed action space; Jev SENSOR; policy BLOCK/ASK/ALLOW. Distinct from pre-exec toolgate. Human is the only kill trigger; identity re-check; shields override; mapped explanations not raw model prose. **Permission vs probability:** operator owns auto-approve thresholds; plugin never self-tunes the safety bar; host deny stays above (omp-greenlight; not a sandbox). **Spoken confirm ≠ auth:** jev-voice-browser `destructive` spoken "confirm" is convenience not a guarantee (control-port reach is the grant; rh-guard owns the gate cousin). **Privilege ≠ verdict:** effect-based shell gate; fast structural rules then Jev Choice + independent risk Nouls; fail-closed (construct-auto-classifier; 0 dangerous / 975 *theirs*). **Jev supplies evidence, code owns authority** (actiongate-jev; a positive score never overrides a deterministic security failure). **Turnstile:** policy first, Jev after permit, missing Jev → Review, replay thresholds. **SEAL:** no seal, no advance; coverage.path visible; mint ≠ product brain. **Never confidently wrong:** TLA+ kernel may escalate, must not return a confident wrong (jev-labs). **Landed-script trust / headless≠auto-approve:** construct (byte-identical to remote default branch; headless escalation is deny-and-report, not auto-approve). **Typed baton:** escalate/continue/abort; gate `allow` never grants (jev-handoff). **Persist constraints across compaction:** conversational policy as structured state; Jev never writes policy; fail-open (pi-heed 98.5%/0 false block *theirs*). **Pi control-plane tool gate / GUI:** pi-jev-control (deterministic fast-path then Jev; GUI < threshold → unknown, never force-click). **jev-use PreToolUse gate:** deny/ask, fail-open, 12/12 *theirs*; Vercel margin fallback. **Commit-msg hook:** commitjev blocks only on a warning; instrument failure is not a refuse. Toolbelt notes: jev-security-scan / jev-decisions / TeoMastro; unofficial jev-cli not ready (≠ jevql); actiongate slogan already §64; **rh-guard owns reward-hack**. **Wrap-as-execution ALLOW/ASK/DENY:** the wrap *is* the tool function; rules first; ASK throws; fail-closed (AgentGhost; **≠** actiongate **≠** toolgate **≠** jev-use; rh-guard owns the gate cousin). **Tool-risk as placement (not a wrap card):** Akshay pedagogy cites LangChain middleware; AgentGhost owns wrap-as-execution. **Human every action:** Essentiel-Jev never authority. **Hourly 1241:** never-execute list; tab-bouncer pinned/audio/current never closed; ORIGIN pause-if-no-Jev. **Hourly 1347:** typed-gate band [0.40,0.60] is refusal; pi-jev-gate fail-closed; choice is the verdict; jev-calibration-arena never acts. **SIGNAL §93:** Cache hit ≠ correctness; score never auto-accepts (rh-guard thin). **SIGNAL §94:** guidance ≠ hook; unofficial ≠ TypeSafe; hosted bootstrap ≠ silent TypeSafe; Nemotron ≠ TypeSafe Jev; not a calibrated replacement (rh-guard thin). **Hourly 1541:** jevguard calibrator/cache/escape; jev-ci-selector CI shadow mode; one-dollar-tahoe TypeSafe Jev defense eval. **Hourly 1740:** Decision Graph Protocol frame→assess→commit; app retains permissions/effects; Jev-first assessor-neutral; guarded commit / receipt/next frame; assessment batching; hard-gating DGP as safety theater; numerous-com/dgp ≠ TypeSafe official. **Hourly 1843:** YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human; default on_unsure=raise is app policy not a System One hard gate; cost_human is VOI not a grant. **Hourly 1943:** memory leases ended by new evidence; six Nouls then fixed rules in code; 0 of 157 false invalidations; unsure → review queue; host keeps the store **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/applied-mappings.md#7-capability-kernel--human-confirmed-gate` |
| Decide → policy → LLM leftover | Typed decide backends share one Answer schema; policy in code routes auto/review/llm; generator writes leftover text only. Noul 0.5 = cannot-tell, never rounded. Score conf 0.0 = flat, never acted on. Hard flags always review. **Inbox cousin:** mailordinal (typed signals → deterministic priority; no leftover LLM required). **Public decide-backend:** classifier.dev (HTTP classification API; leftover LLM is fallback when Jev is down). **Open NAR cousin:** NandhaKishorM/laya (self-hosted Choice/Score/Noul; Router picks ckpt; policy still in the caller). **Compile leftover:** byenzyme/enzyme (Decisions uses Jev; catalyst generation uses a separate LLM; guidance ≠ hook; hosted bootstrap ≠ silent TypeSafe). **Read-only Gmail trays** (jevmail leftover); **macOS proposed writes** (mailjay). **Hourly 1241:** ORIGIN pause-if-no-Jev; validResponse sums-to-1. **Hourly 1347:** jev-calibration-arena never acts; codaaiteam/jev-skill jevtypesafeai.com ≠ TypeSafe. **Hourly 1441:** ha-conversation-jev Jev→Grok; dsh-jev can only gate. **Hourly 1541:** keep/shadow/hybrid/reject; difficulty + policy thresholds + JSONL trace; jev-codex-pilot model + reasoning depth | `references/applied-mappings.md#8-decide--policy--llm-leftover-cascade` |
| Closed-vote computer-use | Code builds every option from observation/goal/facts; decision model only picks; code acts/verifies/undoes. No planner LLM. Host-owned product cousin: app retains handlers/permissions; Jev over live typed actions. **Hot-click cousin:** indexed viewport table → operation+target; code owns the loop; generator only for type (ego-jev). **Adversarial cousin:** Playwright executes, Jev chooses explore/continue (browser-jev; sample not argmax). **Decider≠executor cousin:** Jev picks next-tool/progress/risk/done; LLM only fills args (jeffrey; risk≥0.5 pause; stuck ladder). **Hand no-text steps:** jev-use (plugin; writing stays with the LLM). **Never free-generates:** jev-gpt (WordNet tree of Choices; architecture demo). **OCR+AX productized cousin:** typesafe-computer-use (macOS; exclusive kinds; split questions; post-type Noul still soft; 155× *theirs* one screenshot). **ASR voice-browser cousin:** jev-voice-browser (Playwright; partial-speech wait; spoken confirm ≠ auth). **CU source-study pointer:** dairui1/jev-lab; do not re-card jev-desktop. | `references/applied-mappings.md#9-closed-vote-computer-use` |
| Agent preference lint / semantic gates | Soft project rules as criteria; linter owns hard rules; one Score per rule on the diff; bands + fail-open; name the observation window (edit vs turn). Named Choice escape when conflict ≠ ignorance (typed-evaluation-collapse). Sentence-as-rule cousin: ast-grep `rule:` × Jev `ask:` (mizchi/jev-lint is mizchi/jevlint rename; ~1 in 5 findings wrong *this README*; 13/15 1.00/1.00 older corpus; fail-open no-verdict). Independent of huntedman/JevLint | `references/mixed-architecture.md#preference-lint-and-gates` |
| Dual orchestration (Jev ∩ LLM ∩ MCP) | Jev-as-tool vs Jev-as-outer-loop; schemas are exact state | `references/mixed-architecture.md#dual-orchestration-jev--llm--mcp` |
| "It's just classification" / "not probabilistic programming" / stack-replacement FAQ | Typed judgment is a software primitive, not a new task; marginals are not a joint; Jev is not the only model. **Public pedagogy:** Akshay “Jev Clearly Explained” — LLM hammer; schema-safe ≠ correct; 200×/400× TypeSafe ceiling. **Meaning-grep:** jev-semgrep ≠ semgrep.dev; proposition≠embedding; not a gate. **Hourly 1047:** Jev never authors UI text; jevtest 0.85 still soft / ambiguous band; product bakeoff ≠ architecture duel; majority floor; HA remains execution; compaction-pi ≠ compact ≠ compaction; jevloop mock ≠ quality; Cerebellum `/v1/decide` ≠ TypeSafe; laya-grounded not drop-in. **Queued SIGNALs:** Is GestaltLabs/Jeff-1 logan-markewich/jeff?; Empty stanley findings = the change is fine?; Is findme JevFind?; Swap the conversation model per turn? **Hourly 1144:** Is `.feels()` a new language?; Is apa-agent-harness jev-harness?; Quote grok-bot-jev 13.0× as token savings?; Can Essentiel Jev send mail?; Is enzo-mcp jev-sift?; Treat pigeonhole OTHER as a move?; Is the HF playground live Jev?; Does jev-reliability measure accuracy?; Is clduab11/jev-test the jevtest matcher?; Did jev-rag-benchmark show Jev wins?; Is dairui1/jev-lab BrendanH18/jev-lab?; Is jevmail mailordinal?. **Hourly 1241:** Is ZHUBoer/ego-jev jiangkoumo?; Are jsort scores frequencies?; Did groundedness-judge-bench show Jev wins quality?; Quote jev_playground 83%?; Copy `jev-latest` on Zen?; Does the techstack classifier generate a stack?; Is s1_ruby hunch?; Is judgement jevql?; Is the Rust community SDK official?; Is tpellet/hunch carldaws/hunch?; Quote file-search 15 matches as recall?; Treat linkmap referee as gold?; Is jev-mail jevmail?; Does ORIGIN’s LLM decide?; Gate crawlers on raw `bug_likely`?; Is jevbrain TypeSafe Jev?. **Hourly 1347:** Is judgekit JudgeBench?; Is decide 0.8 80% accuracy?; Does the calibration arena act?; Is openrouter-jev-mcp TypeSafe first-party?; Is FrancoisChastel/jev-code stanley npm?; Is hermes-switchyard Agnes?; Is system-one-skills a judge?; Is typed-gate 0.51 a yes?; Is pi-jev-gate fail-open?. **Hourly 1441:** Paste Foq 100%/ECE as a class ceiling?; Quote rev 32.4 as a bench?; Is robfrase/jev running code?; Treat 27/27 as Harbor?; Is safe-sh pre-exec?; Skip Confirm?; Quote Reranker 0.1667 as live Jev?; Fail-closed sessionwise?; Treat jev-search scores as truth?; Collapse savka777/jev-search into kazuhideoki/jev-search?; Is 400 ms an SLA?; Let the scheduler place Pods?; Screenshot Android for the decision?; Is jevcu closed-vote?; Is ha-conversation-jev HA-Jev?; Can dsh-jev widen tools?; Paste $0.46 as a measured run?; Auto-page from luna-pagerduty?; Quote akpsahan vs-Jev as new / is it Archer?; Let Jev own chess/AV/Post/API Victory?; Is typesafe-go official?; Treat a ledger HIT as correctness?; Collapse hyperspaceai/jevcache into kushals256/jevcache?; Auto-accept a GEPA proposal because the training score rose?; Is sutro-sh/jev-align caiovicentino/jev-align?; Hard-gate enzyme `when asked` / catalyst similarity?; Are catalysts summaries?; Is unofficial JA ModernBERT TypeSafe?; Is Argos1111/jev_local us/jev-local / kunchenguid/local-jev?; Is LFM default the JA ModernBERT backend?; Treat enzyme hosted bootstrap as TypeSafe / silent FALLBACK?; Is Nemotron_Jev a calibrated Jev replacement?; Is djev-dev djev-spark / localjev?; Paste Laya essay vs-Jev as a new bake-off?; Treat Khmer 0.952 conf as competence?. **Hourly 1541:** Does orchestrator pick an LLM by difficulty?; Paste 0.95 as Harbor τ?; Is jev-codex-pilot a bake-off?; Is keep a failed migration?; Summarize quarry pages?; Fail-closed quarry?; Is Frank-ZY-Dou walidboulanouar?; Invent one-dollar-tahoe ASR/FPR?; Is jevguard jevcache?; Skip CI from skip_below 0.05?; Is llama-jev a Noul?. **Hourly 1740:** Is numerous-com/dgp official TypeSafe?; Does assessment p grant an effect?; Is a receipt proof the decision was right?; Hard-gate DGP as safety theater?; Is can1357/jegrep Bentlybro/jevgrep / uehaj/jev-semgrep?; Paste jevgrep 79%?; Hard-gate 0.4/0.2 as concept absent?; Paste $0.01–0.03 as a class ceiling?; Treat OpenRouter/TypeSafe auto-failover as the same Noul?; Is kev OOD 0.76 Jev-equivalent?; Does isolation 4e-6 prove kev = Jev?; Is `/v1/systemone` wire a Noul?; Did Archer land?; Are gut/judge new class-table species?; Hard-gate cost-derived 0.038 as a proof?; Is Kungie/gut tpellet/hunch?; Is Illusion47586/judge judgekit?; Is jev-forge a sixth species?; Clone jev-forge weights?; Merge von Needle 52.6% with 93% n=78?; Is sub-15ms the 62 ms table?; Does von guarantee calibration?; Did Jev win guardrailing?; Is ishaannk/llm-vs-jev a Harbor taskset?; Does Augustus own the steerability fold?. **Hourly 1943:** Is southpolesteve/probably a hunch/feelings/gut/Judge overlay?; Is it tidymodels/probably?; Does the hosted playground run live Jev?; Hard-gate `feels` 80% as a safety proof?; Treat jev-lint rename as a second product?; Paste 17/18 as Harbor?; Hard-gate 0 of 157 as a proof?; Treat boolean @ 0.5 as safety?; Claim 10×?; Treat snapjudge softmax as a Noul?; Treat JevPi 62 tests as quality? **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/faq.md` |
| Feature engineering / multi-criteria analysis | Nouls + Score distributions as named features, weights in code | `references/mappings.md#1-semantic-judgments--features-and-explicit-utility` |
| Selective classification / decision theory | Thresholds from action costs, abstention paths. Train a domain specialist when downstream code **reads the probability**; few-shot hosted API when only **argmax** matters (calibration/VOI, not an accuracy bake-off). **Active-learning triage:** high conf accept / middling expensive teacher / low-or-boundary human; log full distributions. **Do not distill Jev as teacher of record** (~68% ceiling compounds errors; real outcomes stay the targets). **SIGNAL §93:** human labels only; production capture flywheel; score never auto-accepts | `references/mappings.md#2-probabilistic-judgments--cost-sensitive-decisions` |
| Decision tables / circuits / state machines | Judgment predicates, code owns transitions. Language primitive: Ruby `chance`/`pick`/`rate` as control flow (hunch; English-as-config; fail polarity per action). feelings `.feels()` default 0.5 is Noul-0.5-never-rounded; exhaustive BAML `match`; **≠** hunch **≠** southpolesteve/probably (language primitive, not a `.feels()` overlay). apa-persona-engine SM then leftover LLM. **Hourly 1241:** s1_ruby collapse late; `undecided?` abstain; tpellet/hunch exit 3. **Hourly 1347:** typed-judge-kit verdict-in-code; judgekit YAML classify/score/route/verify | `references/mappings.md#3-semantic-predicates--decision-circuits` |
| Retrieve + expensive relevance fn | Bounded rerank of a retrieved shortlist. Decision-native RAG: retrieve wide → decide explicitly → evidence set → conflict resolve → reason only over kept evidence (embeddings stay candidate generators; no universal benchmark). Classify-first MCP: same sandwich on agent I/O (path/url/text → judge; main LLM opens survivors). Meaning-search without embeddings (packed parallel relevance; two-stage outline→zoom). **Meaning-grep** AND/OR/NOT over *thresholded* line Nouls; proposition≠embedding; not a gate (jev-semgrep §86). **Evidence-packet explorer:** index-once, BM25 shortlist, Jev ranks, citable source_of_truth (jevex). Measured pointwise rerank vs a generative reranker (one-run; name the no-RAG arm). **RAG rerank harness:** jev-rag-benchmark; “Jev wins” is not an assumption. **Hourly 1241:** jsort scores are relative; Noul not Choice for scale; groundedness-judge-bench native vs schema-guided; implicit_true included in yes. **Hourly 1347:** nanoprune 2.8MB ECE 2.58%; alsoleg89/decide packing VOI. **Hourly 1541:** quarry evidence projection. **Hourly 1740:** jegrep calibrated path+range Nouls; no embeddings/index/daemon; ~$0.01–0.03 typical; agent --json; can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep. OpenRouter/TypeSafe auto-failover is silent FALLBACK. **Hourly 1843:** variable-N option scoring as the trainable object (dynamic candidate bags not fixed label sets). **Hourly 1943:** retrieve by relevance not resemblance; samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/mappings.md#4-retrieval--bounded-semantic-reranking` (independent TREC DL2019 benchmark: Jev zero-shot best MAP 0.4748, nDCG@10 0.683 vs tuned monoBERT 0.718 — competitive, not dominant) |
| Store as semantic index (SQL / SQLite / zoxide / dataframe) | Cheap exact predicates first; typed questions on the remainder. In-engine extension (sqlite-jev / pg-jev) vs **judgment outside the store** (jevql CLI; vanilla Postgres never sees `jev()`) vs path index (joxide) vs dataframe columns (jevpandas / jevframe). **ORDER BY over probs is a ranking job:** calibration ≠ sortable; measure pairwise inversion / Score ordinality / two-decimal ties (jev-orderby-bench) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` |
| Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident; compaction mutations/shell operators → keep_full; stdout prune: ≤10k/JSON-diff-whole-doc untouched, then Noul; Cua-S1: plan≠execute, dry-run, fail-closed checkbox/fill; Stagehand extract: schema/completion-gate/screenshot-always-LLM then pick, else LLM; skills→oxlint: AST/precheck prove, guidance whole-file in state, remainder Noul — not a hard gate; pre-exec toolgate: allow/block/review — Jev is not authorization; guard error/timeout stops; distinct from capability kernel interlock — secrets never in agent, closed action space. OMP prompt suppression: operator owns thresholds; plugin never self-tunes; host `bash.patterns: deny` stays the floor (omp-greenlight — not a sandbox). Human-confirmed kill: Jev recommends, human is the only trigger, identity re-check, shields override, mapped explanations (port-cleanup). Effect-based shell: fast-allow/deny prove, then Jev remainder; privilege ≠ verdict (construct-auto-classifier). Wrap-as-execution: rules prove allow/deny, Jev remainder, ASK throws, fail-closed (AgentGhost). SLO router: exactness raises quality floor, never overrides capability; Jev features fail-open to local (slo-router). **Hourly 1740 DGP:** JSON Schema proves record structure; app owns freshness/authorization/idempotency; assessment is the remainder sensor; commit fail-closed (numerous-com/dgp; hard-gating DGP as safety theater). **Hourly 1843:** thresholds derived from costs not hard-coded; the cost table is policy; the Noul is a SENSOR; hard-gating 0.038 is theater). **Hourly 1943:** noul_threshold 0.5 decoder not a proof; softmax over allowed tokens ≠ Noul; no shipped rule has severity error **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` |
| Value of information / gather as an act | Pay for another observation only if EV(decision) improves more than cost; abstain from calling *any* model when a regex already answers (meta-VOI). Fail-open wake/resume: skip the LLM turn only if the judge answers and p(wake) is low (Horvitz). Selective memory: verbatim ledger + scored recall (carryforward; rules never judged; 9×3 is a hint). Classify-first read: pay for a full agent open iff relevance (or typed question) says it might change the act (jev-sift; errors/truncation ≠ irrelevant). **Training-data VOI:** spend expensive teacher/human labels only where confidence says they change the outcome (jev-triage); do not distill Jev as teacher of record. **Decision-model latency cost:** measure p95 of sync Jev on the hot path before claiming routing (slo-router 77.93→490.38 same routes). **Human-review VOI:** attention filter that never blocks the agent and never says green unless sure (jev-lens). **Skill-library VOI:** load a skill iff it changes the next step; abstention first-class (skillranker). **tools≠use:** an MCP memory tool the agent never calls is not VOI (carryforward 0/4; SessionStart hook). **Pre-send token-econ:** compress before first send (dizk/jev-lens; post-send prune cost 17% more *theirs*). **Same-intent cache admit:** skip the LLM iff Jev says same intent (jevcache 0 FP/100 *theirs*; fail-open). **Human-feed VOI:** worth-your-attention before click (ThinkyMiner/Winnow). **Second-call VOI:** skip the narrating main-model turn (hermes-jev-router WHETHER/HOW/WHAT). **Hunk-review VOI:** score PR hunks before the generative reviewer (prune-review). **Intent-search VOI:** pay to judge places the diff did not touch (jev-intent-review). **Empty compact-proxy skip:** IPECTER context-pruner **and** jev-runway slogan only. **Batch ranking VOI:** one request per ten (jevfeed). **No-text-step VOI:** hand the judgment, keep writing on the LLM (jev-use 186 vs 2,672 ms *theirs*). **Files-to-read VOI:** jevex n=16 160s→69s *theirs*. **Commit-attention VOI:** commitjev middle band never rounded. **Pi compact VOI:** pi-jev-compact keep/drop vs LLM summary. **Escalate-under-threshold VOI:** classifier.dev smart re-asks single-label <0.7; multi-label ignores (re-judge worse). **Escalate-without-stall VOI:** khordoo S2 is async one-use; S1 never waits; log consumed not arrived. **Split-question CU VOI:** kind/item/site/offscreen in one request; exclusive actions; perception rebuild is the expensive gather (typesafe-computer-use; 155× *theirs* one screenshot). **Partial-speech VOI:** complete Noul; closed-set may fire; free-text waits (jev-voice-browser). **Evidence-synthesis two-pass VOI:** choxos/jev-reviewer fan-out then absolute Noul; *Not found* cheaper than paraphrase (≠ egma-ai). **Script-before-p VOI:** NandhaKishorM/laya Router routes before the forward pass because Khmer 0.000@0.952 gating cannot catch. **Path-then-window VOI:** JevFind scores paths first (0.25) then windows (0.55); `--file-threshold 0` when recall matters. **Frontier cascade VOI:** jev-frontier-bench pay Fable only when Jev top p < 0.9 — cut is **in-sample**. **Compaction-pi VOI:** fail-open to built-in LLM summary. **Beam-search FS VOI** (findme). **Cache-vs-worker VOI** (jevsubrouter). **Cheap decision-layer VOI** (grok-bot-jev). **Question-preflight VOI** (jev-reliability). **Hourly 1241:** jev-file-search scores not calibrated accuracy; jev-linkmap Jev never sees S2 prose. **Hourly 1347:** alsoleg89/decide packing VOI; 0.8 ≠ 80% accuracy; Jev-Calibration Platt ECE 0.117→0.052. **SIGNAL §93:** fingerprint after redact; recall vs decide; publish fingerprints+answers; VOI of cache hit; hyperspaceai/jevcache ≠ kushals256/jevcache; human labels only (uncertainty acquisition). **Compile-time index VOI:** pay for Jev at compile to shape catalysts, then retrieve through questions (byenzyme/enzyme; catalysts ≠ summaries; guidance ≠ hook; hosted bootstrap ≠ silent TypeSafe). **Hourly 1541:** quarry evidence projection. **Hourly 1740:** assessment batching; ~$0.01–0.03 typical; no embeddings/index/daemon; agent --json . **Hourly 1843:** cost_human is a gather act; auto-batching same-object questions is measurement economics; YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human. **Hourly 1943:** 133★ / forks 10 live; build calibrated classifiers from human feedback; one calibrated yes/no per memory in one request **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/mappings.md#6-value-of-information--gather-as-an-enumerated-act` (**Hypothesis** until a labeled act/outcome log; 149-row receipt is Empirical as a shape; wakegate 21/21 is smoke; jev-triage is Empirical as README architecture; slo-router live analysis is Empirical as a *negative* on sync Jev; jev-lens is Empirical as README architecture) |
| Signal detection / ROC | Criterion and operating point from costs and base rate, not accuracy. Operator owns the criterion; a plugin must not self-tune the safety bar (omp-greenlight 40.9% / 0 of 94 *theirs*). Exactness raises a quality floor — it must not override capability/context (slo-router). Privilege ≠ verdict (construct; `sudo status` can be safe). **Ranking ≠ calibration:** never hard-threshold raw p as a frequency (does-jev-confidence; AUC ~0.91, stated ~75% vs human ~10%). **Hourly 1347:** Jev-Calibration Platt ECE 0.117→0.052; typed-gate band [0.40,0.60] is refusal; 0.8 ≠ 80% accuracy. **OOD / AUC ≠ ECE:** sign of miscalibration flips by type (jev-ood-calibration; unknowable policy label mean p 0.74). **Sureness over vectors:** max_prob/margin/entropy/gini/perplexity → CERTAIN|…|CLUELESS; Choice confidence = max_prob (most generous) (how-sure-is-jev). **Conflict ≠ ignorance:** Noul collapses both; named Choice separates (typed-evaluation-collapse). **BBQ stereotype/uncertainty:** 97.28% / bias 0.04/0.34 *theirs*; not a general bias cert (jev-bbq-experiment). **Cookbook moderation dial:** 5 Nouls; samples not benches (jev-cookbook). **Dual-channel ECE / like-for-like:** openJev-verdict-2.0 claim-audit (correctness-head ≠ distribution ECE; PR #1). **Commit middle band:** commitjev never rounds 0.35–0.65. **English-checkpoint OOD:** laya-multilingual (Khmer 0.000@0.952; gating cannot catch). **Coverage ≠ correctness:** chakuho 8B coverage 1.00 while `__none__` collapses. **Peaked schema-scorer:** rankings not calibrated confidences. **Escalate-under-threshold:** classifier.dev 0.7 is *theirs*; multi-label does not share it. **Laya 0.85 gating still soft:** README recipe, not Harbor-calibrated; post-T ECE ≠ raw ECE; Banking77 token-budget (NandhaKishorM/laya). **Majority floor / calibration ≠ discrimination:** job-posting-triage floor 0.947; Jev on the floor; llm_local ECE 0.947. ChaosNLI JS: uniform can beat Jev (`notes.md` §87). **Question preflight as SDT** (jev-reliability). **Triage buckets as criterion** (dairui1/jev-lab). **Hourly 1241:** Noul not Choice for scale; confidence ≠ winner p; jsort scores are relative. **Hourly 1740:** jegrep calibrated path+range Nouls; ranking fail-open; 0.4/0.2 still soft; kev family OOD 0.76–0.77 vs Jev 0.86; replica honesty. **Hourly 1843:** typed judgments vs chat judges on guardrailing; ishaannk/llm-vs-jev cross-note only; deeper integrity fold is rh-guard; nothing wins outright; can be argued out of guarding; thresholds derived from costs not hard-coded. **Hourly 1943:** embedding resemblance misses the allergy; pointer mode 17/18 19/20 *theirs*; ~1 in 5 findings wrong *theirs* **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/mappings.md#7-signal-detection--criterion-not-accuracy` (**Hypothesis** for non-SWE plots) |
| Org / safety control structure | Sensor ≠ constraint (Leveson); STPA if the sensor lies. Capability kernel: Jev SENSOR, policy.py constraint; secrets never in agent (interlock). Host deny fires ahead of Jev prompt-suppression (omp-greenlight). Judgment ≠ permission: Jev never grants skill access (skill-broker outline). Contracts on effects, not tokens (construct). Attention filter ≠ permission gate (jev-lens). **Jev supplies evidence, code owns authority** (actiongate-jev; positive score never overrides a deterministic failure). **SEAL coverage ledger** (exception queue visible; mint ≠ product brain). **Never confidently wrong** (jev-labs; escalate instead of hard-gate). **Physical-world S1:** typed answers as HA sensors; not for locks/heaters (HA-Jev). **Gate never grants** (jev-handoff). **Conversational constraint sensor** (pi-heed; Jev never writes policy). **Pi control-plane sensors** (pi-jev-control: router/gate/retry/sieve/review/click; GUI never force-click). **jev-use gate never grants** (fail-open PreToolUse). **Inbox policy in code** (mailordinal; model never sets queue order). **Agnes branded as Jev is not a sensor** (hermes-plugin-jev). Toolbelt sensors ≠ policy (jev-security-scan / jev-decisions); unofficial jev-cli not ready; actiongate slogan §64; rh-guard owns reward-hack. **Wrap-as-execution:** wrap *is* the actuator (AgentGhost; ASK throws; fail-closed; rh-guard owns the gate cousin). **Meaning-grep is not a gate** (jev-semgrep ranking fail-open; rh-guard skip). **HA remains execution:** ha-switchboard (Jev SENSOR; HA constraint + actuator; one bounded LLM handoff; **≠** HA-Jev). **Authorship / jevtest-as-merge-seal:** rh-guard owns the gate cousin. **Empty findings ≠ approval** (stanley-code; `notChecked` coverage ledger). **Human every action** (Essentiel-Jev never authority). **Atom then sense** (enzo-mcp UNKNOWN). **Hourly 1241:** ORIGIN pause-if-no-Jev; jev-crawlers risk bands never raw boolean; jevbrain AUTO_ACT is not a Noul. **Hourly 1347:** jev-calibration-arena never acts; typed-gate band [0.40,0.60] is refusal; pi-jev-gate fail-closed; choice is the verdict. **Hourly 1541:** jevguard calibrator/cache/escape; jev-ci-selector CI shadow mode; one-dollar-tahoe TypeSafe Jev defense eval. **Hourly 1740:** Decision Graph Protocol frame→assess→commit; app retains permissions/effects; Jev-first assessor-neutral; guarded commit / receipt/next frame; hard-gating DGP as safety theater; numerous-com/dgp ≠ TypeSafe official. **Hourly 1843:** cost table is policy; Noul is SENSOR; default UNSURE raises; deeper integrity fold is rh-guard. **Hourly 1943:** questions/plans/directives are not evidence; host keeps the store; six Nouls then fixed rules in code **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/mappings.md#8-control-structure--sensor--constraint-leveson` |
| Search / control loops (any domain) | Algorithm stays yours; judgment substitutes one classifier step. **Decider≠executor:** Jev picks next-tool/progress/risk/done; LLM only fills args (jeffrey; pick ≠ fill; still not a planner-writer). **Tree-of-Choices writer:** jev-gpt never free-generates (one question per word). **Pick≠write plugin:** jev-use. **1-token selector:** chakuho reads next-token mass over declared labels (not a writer). Robotics text-state (geometry-as-text, not pixels); khordoo: no graphical input; S1 flies, S2 advises without stalling; physics owns collisions; do not replace A* / a solver with a Noul. **OCR+AX desktop CU:** typesafe-computer-use (hosted Jev; never screenshot-to-frontier for the decision; exclusive actions; split kind/item/site). **ASR voice-browser:** jev-voice-browser (Playwright executes; partial-speech wait; spoken confirm ≠ auth). **Hybrid S1:** anima3 closed verb menu + hard safety first; Qwen logprob default; jeff confidently flat on magnitude; scene is a11y tree not screenshot. **Full-distribution optimizer:** jevloop UCB1+CEM; no LLM in the loop; mock default. **NL memory → beam-search FS** (findme; the *algorithm* is beam search). **Persona state-machine** (apa-persona-engine). **Hourly 1241:** ORIGIN pause-if-no-Jev; validResponse sums-to-1; jev-crawlers risk bands never raw boolean. **Hourly 1347:** smartdio/jev-browser-agent ≠ ZHUBoer/ego-jev; Dakai/omp-jev-web DONE ≠ proof; hari007sh/jev ≠ dannote/jev. **Hourly 1441:** droidjev screenshot-free; Tewoto1 jevcu planner still writes (not closed-vote). **Hourly 1541:** Frank-ZY-Dou/awesome-jev robotics/3D/control; difficulty + policy thresholds + JSONL trace. **Hourly 1740:** Decision Graph Protocol frame→assess→commit; guarded commit / receipt/next frame; assessment batching; Archer-arch fidelity; replica honesty. **Hourly 1843:** cost-sensitive decision theory × System One probabilities → control flow; judgment vs generation; deterministic execution after probabilistic judgment; exactly one app-owned callback; explicit uncertain branch. **Hourly 1943:** Jev IS the if-statement; Jev-first Pi agent loop; slow-LLM fallback; explicit action menu / CandidateSource unimplemented; 62 tests wiring not quality; direwolfiy/JevPi ≠ standardagents/jevpilot ≠ pi-jev-control **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/mappings.md#9-search--control-loops--one-substituted-classifier-step` |
| Spec property pipeline | Rank candidate props; checker owns validity | `references/mappings.md#10-spec-property-pipeline-hypothesis` (**Hypothesis**) |
| Alloy instance loop | Cluster CEXs; Analyzer owns in-scope truth | `references/mappings.md#11-alloy-instance-loop-hypothesis` (**Hypothesis**) |
| Runtime assurance sandwich | Abstain → RV/monitor → act | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis` (**Hypothesis**) |
| DST multiverse triage | Cluster failing seeds/timelines; regress on the same seed | `references/mappings.md#13-dst-multiverse-triage-hypothesis` (**Hypothesis**) |
| Durable agent control | Resonate protocol settles promises; Jev gates inside a step | `references/mappings.md#14-durable-agent-control-hypothesis` (**Hypothesis**) |
| Assignment hybrid | Soft affinity + hard solver. **Empirical as shape:** slo-router constrained min-cost s.t. quality+SLO floors; Jev features, not the sole gate (p95 77.93→490.38 same routes *theirs*) | `references/mappings.md#15-assignment-hybrid--soft-affinity--hard-solver-hypothesis` (**Hypothesis**; slo-router Empirical as *shape*) |
| Situated density (Shirky) | Aggressive soft loops only inside a named community | `references/mappings.md#16-situated-density-shirky-hypothesis` (**Hypothesis**) |
| Input brittleness / paraphrase stability | Synonymous wording that swings p → abstain or rewrite. **Preregistered framing measurement** (jev-reliability) | `references/mappings.md#17-input-brittleness--sensitivity-calibration-selective-abstention-hypothesis` (**Hypothesis**) |
| Structural prove ∩ soft remainder | Allowlist *proves* the easy verbs; judge only unlisted leftovers; fail-open (cannot block). Effect-based cousin: fast-allow/deny <1ms then Jev remainder, **fail-closed** on execution (construct-auto-classifier). **Wrap cousin:** rules first then Jev remainder, ASK throws, **fail-closed** on execution (AgentGhost) | `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` (**Hypothesis**; jevgate/OCR shapes Empirical; construct Empirical as certification; AgentGhost Empirical as README wrap) |
| Effect-oriented state-machine loops | Soft predicates on transitions; code owns the transition | `references/mappings.md#19-effect-oriented-state-machine-loops-hypothesis` (**Hypothesis**; ZIO client, not Effect.ts) |
| Agent self-supervision / on-track detection | Pre-gate → output judge → done-check → supervisor nouls. S1 reflex keeps control; optional S2 is one-use advice (khordoo delta: never stall; log consumption; Local ≠ localjev; 20% still soft). **OCR+AX desktop CU:** typesafe-computer-use (never screenshot-to-frontier for the decision; `done` ≠ success; 0.4/0.5 still soft). **ASR voice-browser:** jev-voice-browser (never waveform-to-Jev; spoken confirm ≠ auth). Claim/evidence Stop (anti-hallucinated-done); bounded Pi supervisor (shadow recovery, never generates commands). OMP/pi `jev_acceptance_gate` + `jev_route` fail-open (confidence: 0 on missing Jev). OMP prompt suppression: operator owns bar; host deny stays above (omp-greenlight). Closed-vote CU: no planner LLM (JevOnly); host-owned product (waymode). Effect-based pre-gate fail-closed (construct-auto-classifier). Stop-hook attention filter: never blocks the agent; never green unless sure (jev-lens). Runtime authorize: Jev supplies evidence, code owns authority (actiongate-jev; fail-closed on deterministic security failure). Turnstile clone: observe/enforce + replay; Jev never grants what policy denied. SEAL: generation fills Candidates; only a Seal advances. jev-labs: quorum+stability; escalate when evidence degrades. jev-handoff: typed escalate/continue/abort; inverted loop; fail-open. hermes-jev-router: WHETHER/HOW/WHAT; skip next main-model when evidence is enough. browser-jev: Playwright executes, Jev chooses; sample-from-distribution. jeffrey: Jev→tool→Jev; LLM fills args. pi-heed: persist user constraints across compaction; check side-effecting calls. jev-decisions: reviews are advice, never stop commands. **jev-use:** fail-open PreToolUse gate; Vercel margin fallback. **pi-jev-control:** control-plane sensors; GUI unknown never force-click. **Wrap-as-execution:** ALLOW/ASK/DENY on the actuator path; ASK throws; fail-closed (AgentGhost; **≠** jev-use fail-open; rh-guard owns the gate cousin). **Hybrid S1:** anima3 hard safety first then pick among valid verbs; low conf → the rule. **Decision-as-assert:** jevtest meaning Noul; ambiguous band never rounded; 0.85 still soft. Empty findings is **not** a self-supervision pass (stanley-code). Shadow then honor (not a merge seal): apa-agent-harness. **Hourly 1241:** ORIGIN pause-if-no-Jev; jevbrain AUTO_ACT is not a Noul; runWorkflow completed ≠ success. **Hourly 1347:** jev-calibration-arena never acts; typed-gate band [0.40,0.60] is refusal; Dakai/omp-jev-web DONE ≠ proof. **Hourly 1441:** typesafe_agent_gates 27/27 / 31/31; pastepilot Confirm before act; droidjev screenshot-free; Tewoto1 jevcu planner still writes; dsh-jev can only gate; ha-conversation-jev Jev→Grok. **SIGNAL §93:** production capture flywheel; human labels only; score never auto-accepts. **SIGNAL §94:** guidance ≠ hook; unofficial ≠ TypeSafe. **Hourly 1541:** jevguard calibrator/cache/escape; jev-ci-selector CI shadow mode. **Hourly 1740:** Decision Graph Protocol frame→assess→commit; app retains permissions/effects; guarded commit / receipt/next frame; hard-gating DGP as safety theater. **Hourly 1843:** exactly one app-owned callback; explicit uncertain branch; default on_unsure=raise is app policy not a System One hard gate. **Hourly 1943:** otherwise maybe / confidence gate; chaos samples after the gate; unsure → review queue **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/agent-self-assessment.md` |
| Optimizer/program frameworks (Ax, DSPy) | Typed fields → one provider request; judge metrics; threshold discipline. Ax and DSPy climb LM-program knobs only. Typed control plane sits *around* DSPy (ontology/security/confidence/state/allow-list); DSPy drafts AFTER route+action. **Full-distribution critic (no LM):** jevloop UCB1+CEM over edit ops in code; mock mode default (control-loop demo, not a Jev quality headline). **GEPA alignment loop:** human labels only; score never auto-accepts; production capture flywheel; sutro-sh/jev-align ≠ caiovicentino/jev-align; GEPA + System One | `references/optimizer-integration.md` |
| Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. Specialist composition stays **Hypothesis**; open multimodal decide (blackwood-rlcd) is a named receipt. Structured observe→decide→verified-act (no screenshots) is a computer-use speed-layer receipt (Jev or GLiNER2 or Cua-S1 specialist; Cua-S1 source-only, not TypeSafe Jev). Stagehand experimental Jev is the same job inside a major harness (pick-and-copy extract; 37/75 no-LLM ~0.5s vs 4.37s is *their* card; pick ≠ replacement). **OCR+AX desktop:** typesafe-computer-use (MIT **427★**; hosted Jev; $0.0002/155× *theirs* one screenshot, not a taskset). **ASR voice-browser:** jev-voice-browser (MIT **103★**; 27/27 fixtures *theirs*). **Laya-class vision:** thaitea/laya-vision-smolvlm-256m (SmolVLM; `score` untrained; CC-BY-NC-SA; **≠** blackwood **≠** Archer). Same section as the row below | `references/validation.md#eval--hill-climb` |
| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score). Harbor-style frozen protocol vs constrained LLMs: DMB (accuracy/calibration/latency/cost; raw logs). Feedstock: jevals-data CC-BY-4.0 boards + JSONL (recompute-from-logs). Collab-arm curriculum: llm_autonomous vs scripted_plus_jev vs llm_plus_jev (Wilson / McNemar). Negative: combinatorial grid assembly ≠ extractive (ARC-AGI Direct Jev 4/400). Harbor on/off routing: chess-engine tasks, hidden perft verifier, one-run preliminary (jev-gateway-bench). Pair CI merge-gate with Harbor + rh-guard. Harbor needle/noise stdout prune: jev-pruner (manual sweep theirs; plugin eval cannot reach Jev → fail-safe original; Terminal-Bench pilot is integration not a full bench). Cua-S1 specialist form: source-only (metric names, no checkpoint scores; not TypeSafe Jev). Stagehand extract pick-and-copy: 37/75 no-LLM ~0.5s vs baseline 4.37s (*their* 25×3; pick is a fast path, not a replacement; draft stack #2951–#2955). Pre-registered independent eval: jev-baselines-eval (**both AMBIGUOUS**; cascade sign-flip at exact parity; confidence=1.0 theater; encoder-with-labels wins; serving-path ≠ model-speed). Healthcare Harbor-shaped: explore-typesafe-ai (synthetic FHIR; not clinically validated). Honest-negative PDF: databricks-jev-pdf-lab (no quality-equivalent Jev payoff). Meaning-search Harbor-shaped: jevgrep 79% top-5 vs BM25 40% / grep 20% on stripped repos (keyword still wins exact strings). Measured RAG rerank one-run: Jev-RAG ≥70% cost / 72% latency vs Spark rerank (full-context Spark still faster). jevals-shaped oxlint: Phoenix fixtures agree with the human answer key. Native-probability calibration arena (jev-arena Brier 0.0059 / ECE 0.0620 *theirs*; overconfident in low bins; fan-out 2 requests). Typed control-plane bake-off metrics (dspy-control-plane; offline stubs ≠ quality). Tetris Jev vs Haiku demo (not a rigorous eval). Fan-out suite: sonar heatmap-as-policy; vickrey CDF then code bids; bracket Brier vs Elo (live trailed Elo). Domain specialist vs few-shot hosted: Domain-jev-maker matched-precision KL / r / McNemar (independent CLINC gold). Cascade compare arms: jav-email-cascade jev vs gen-json vs gen-logprob (mock gen-json flat-confidence is *their mock*). ORDER BY ranking family: jev-orderby-bench six gates; Score ordinal 0.143 weak link; 53-way 0.99 tie; recodelabs batch-40 fails ranking. Class-backend economics: jeff GLiFormer ~$2.6 vs ~$15.6 (~6×) L4 HTTP; A10G direct ~$0.65 (~24×); AG News 75.5% vs 90.5%. Harbor Jev vs local MLX PCD vs AR JSON: system-one-benchmark toxic-chat n=50 Jev 84.0% / Brier 0.1096 vs PCD 52% / Brier 0.3884 (O(1) speed ≠ calibrated Noul). Evidence-packet explorer: 1/8→6/8 SWE-bench Verified finish *theirs* (n=8, empty-as-miss, author-run). Meaning-grep LLM-as-judge: jev-semgrep precision 0.94 / recall 0.98 *theirs* (10×51; not Harbor; dedicated §86; **51★** ephemeral). OMP prompt-suppression: omp-greenlight 1,013 calls / 10 sessions; default **40.9%** prompts removed / **0 of 94** unsafe auto-approvals on labelled corpus (operator owns bar). Eval-instrument hygiene: dinostomp `jev` if-statement tests (accuracy / p(yes) cut / ECE / blank lean / rewording); FINDINGS 189 / 99 against itself. SLO routing latency cost: slo-router p95 **77.93 → 490.38 ms**, same routes/accuracy *theirs* (eight-row demo is not a benchmark). Effect-gate certification: construct-auto-classifier Jev **0** dangerous / 975; every chat model leaked. INSTRUCT_JEV 119-row Choice/Noul/Score instruct seed. **Evidence-gated packs:** jev-packs nine verified packs (accuracy/ECE/cost/latency on pinned jev-1.13; no numbers, no endorsement). **Ranking ≠ calibration:** does-jev-confidence 8,000 human-annotated judgments; AUC ~0.91; stated ~75% vs human ~10%; two-parameter recalibration removes ~96% ECE without changing rank. **Hot-click CU:** ego-jev HN/wiki n=3 medians ~2× vs per-step LLM (high variance; not a bench). **Verbatim compact vs summarize:** jev-compactor 64.5% / 366 ms / $0.0004 / 0 hallucinated paths / 4 of 4 facts vs Sonnet summary 96.2% / 6.1 s / 1 invented path (one synthetic session). **Eval integrity cluster (no leaderboard theater):** atlas receipts (type-safe ≠ correct); jev-frontier-100 Jev 77.0% vs Qwen3.5 4B/2048 96.7% (budget-attached; exploratory); jev-ood-calibration 900 tickets ECE 0.107 = 4.4× floor, Choice/Score T~3.3 vs boolean T 0.66. **JevBench v1.1:** Capability/Speed/Cost → Main Score; calibration reported not scored; native vs verbalized; partial runs not ranked. **Never confidently wrong protocol:** jev-labs 1,080 golden 0 wrong under chaos (escalate; not a proof of zero). **Sureness 60-q:** Choice confidence = max_prob. **Pre-send views:** dizk/jev-lens 79% fewer tokens / 500 SWE-rebench *theirs*. **Product-arm compact:** jev-compactor 73% / 350 ms / 4 of 4 vs shipped summarizers. **tools≠use:** carryforward 0/4. **openvons** LM 0.916 vs 27B 0.875 / JevPick 3.2–4.8× *theirs*. **jevcache** 0 FP / precision 1 / recall 0.38 / n=100 *theirs*. **zeroshot-vs-bert** +0.05–+0.13 vs DeBERTa-c; contamination DiD; ~230 labels *theirs*. **ThinkyMiner/Winnow** 80%/90%. **OpenJev** 45/60. **semif-serve** 1164 vs 178 ms. **pi-jev-skill-bench** harness (43 gold; no live numbers this pass). **typed-evaluation-collapse** Noul band vs Choice p=1.0. **DGUI_HYPERMEM-JEV** 6-row flywheel. **jevassert landed** record/replay CI (accuracy/ECE/Brier/cost/latency; exit 0/1/2; McNemar). **jev-packs pairing** 2,990-case matrix: Jev/Sonnet 5 accuracy tie, Jev better calibrated 7/9, ~250× cheaper *theirs*. **jevarena** failure-finding (≠ jev-arena; harness not findings). **BBQ** 58,492 / 97.28% / $0.3429 *theirs*. **jev-lint** (jevlint rename) 13/15 1.00/1.00 *theirs*. **grande** JGLUE 0.614/0.853 + 270M 0.710/0.710 *theirs*. **local-jev** done 30%/shape 57% vs Jev. **pi-heed** 98.5%/0 false block *theirs*. TeoMastro summary.md 404 this pass; **Harbor SGR-judge contract** (jev-judge-bench; 21 offline tests; canaries ≠ quality; **no quality headline yet**; ≠ jevarena/jevbench); **cookbook samples not benches** (jev-cookbook 16–36; 425/$0.015); **competing NAR claim-audit** (openJev-verdict-2.0 77.10%/0.0636/0.0144 *theirs* unverified; PR #1 throughput≠latency / Laya parity / like-for-like ECE; ≠ IamBusy/OpenJev); **chakuho GUI 336** 27B 95%/92% vs Jev 89%/82% *theirs* (coverage ≠ correctness); **jevinf** 2.57×/2.27× 100% argmax (not ECE); **jevex n=16** 160s→69s / $8.74→$3.13 / 16/16 both arms *theirs*; **commitjev** 13 labelled / 0 false on 5 clean *theirs* (small control); **laya-multilingual** MASSIVE 0.366/0.387 vs 0.227/0.733; ECE 0.314→0.106 after T *theirs*; **schema-scorer** v2 Choice 0.841; **open-jev-laya-bench / jev-tree-choice-cap / INSTRUCT_JEV HF 401** this pass; jevlogs GitHub 404 + HF 401; **classifier.dev vs_jev** tracked JSON (F1 0.887 / 230–232 ms; AG News 87.7%; granite 0.546 vs advertised 0.800 *theirs*; n=7 train-on-test; read eval/README); **githubnext/localjev** 1,200-req prompted-JSON bake-off (Qwen3.6 76.7% / Gemma 26B 75.0% / DiffusionGemma 74.2% short *theirs*; not logits; not calibrated; ≠ kunchenguid/local-jev); **NandhaKishorM/laya vs-Jev** third-party unpublished-here (Banking77 0.425 vs Jev 0.870; post-T ECE 0.081 vs 0.246; 0.766 fine-tune not zero-shot; T4 32.8 ms *theirs*); **external openjev census** (tweet ≠ v1.1 GitHub artifact; watch jev-models; likes ephemeral); **JevBench v1.2** (534 decisions; geo-mean; Jev 75.3 / SemIf 74.6 *theirs*; cal ON; v1.1 87.6 not comparable; option-order 72→21; ×2 latency assumption; many costs est.; Laya absent gap); **OCR+AX desktop CU cost table:** typesafe-computer-use $0.0002 vs Opus $0.032 (155×) *theirs* one screenshot; 0.4/0.5 still soft; **≠** Harbor taskset; **ASR voice-browser fixtures:** jev-voice-browser 27/27 / ~300 ms / ~$0.0002 *theirs* not a Harbor taskset; **wrap-as-execution (not a quality bench):** AgentGhost MIT **2★**; ASK throws; fail-closed; no Harbor numbers; **JP genre atlas (tweet, not scores):** @studio_yebisu stars research-time; not verified evals; likes ephemeral; **external pedagogy (article, not scores):** @akshay_pachaar 200×/400× TypeSafe ceiling; schema-safe ≠ correct; likes ephemeral. **Hourly 1047 Harbor:** jev-frontier-bench Jev 72.5% ECE 0.161 vs Fable 84% ECE 0.064 *theirs* (cascade 82.5% at $4.41/1k in-sample 0.9; ChaosNLI JS Jev 0.149 worse than uniform 0.127; **≠** frontier-100); jev-gliclass-bench Jev 78% vs GLiClass 40% vs majority 49% n=100 *theirs* (product bakeoff; GLiClass flattened); job-posting-triage majority floor 0.947 / tfidf fitted wins / Jev on the floor / llm_local 1000/1000 conf 1.0 ECE 0.947 *theirs*; jev2ui Jobs 11/11 vs Baseline 10/11 valid A2UI *theirs* (valid ≠ good); anima3 offline 20-tick + jeff 0 admitted *theirs*; fast-jev-compaction-pi ~50× *theirs*; jevloop mock default; laya-vision 75.2%/ECE cal 0.034 n=8235 *theirs* (`score` untrained); Cerebellum 94.92% vs Jev 81.1% *theirs* unverified competing NAR — do not endorse; laya-grounded phishing 0.611→0.512 / ECE 0.156 *theirs*. **Open LoRA replica vs hosted Jev (acc vs ECE):** GestaltLabs/Jeff-1 n=9730 acc 0.8183 ECE 0.0807 vs Jev 0.8283/0.0932 *theirs*; **Jev-first bounded agent (coverage, not a bench)** stanley-code; **Sub-agent dispatch counts (not dollars)** jevsubrouter. **Question preflight (Nothing about accuracy)** jev-reliability noul-gate flip 0.0%/12.5%/3.6% *theirs*; **Preregistered hallu bench (bars, not results)** clduab11/jev-test; **RAG rerank harness (no headline)** “Jev wins” is not an assumption; **Triage vs Haiku (synthetic 120)** dairui1/jev-lab urgent 91% vs Haiku 79%; **Grok Bot A/B proxies (not tokens)** grok-bot-jev. **Hourly 1241:** groundedness-judge-bench native vs schema-guided; implicit_true included in yes; jev_playground 0 promotions; routing-backtest 0.0447%; jsort scores are relative; jevbrain AUTO_ACT is not a Noul. **Hourly 1441:** Foq ~25ms/2.2GB local; Jev-Reranker live Jev not yet measured; jev-classification-benchmark specified not run; jev-luna-pagerduty p≥0.50; golergka/jev-plays-starcraft-2 UI-verified ≠ API Victory. **SIGNAL §93:** CI replay as Harbor cousin; Cache hit ≠ correctness; score never auto-accepts. **SIGNAL §94:** JGLUE JNLI 92.62% / JComQA 92.40% *theirs* unofficial JA ModernBERT; Laya essay numbers *theirs*; Gemma ~0.2s *theirs* not Harbor. **Hourly 1541:** one-dollar-tahoe TypeSafe Jev defense eval; llama-jev llama.cpp replica. **Hourly 1639:** OpenCode jev-pruner unit tests ≠ Harbor; do not copy tamaratran 24/24; jev-webagent-bench empty stub. **SIGNAL §97:** GLiNER2 native Apple path; README 0.99 fixture; no Harbor; default threshold 0.1 still soft. **Hourly 1740:** DGP 106 tests ≠ Harbor; mock resolver ≠ Jev; jegrep no published Harbor (do not copy 79%); kev family OOD 0.76–0.77 vs Jev 0.86; replica honesty. **Hourly 1843:** von n=78 / 93.0% *theirs* ≠ Harbor; llm-vs-jev RESULTS generated from summary.json ≠ Harbor; nothing wins outright; competing NAR claims / replica honesty. **Hourly 1943:** pointer mode 17/18 19/20 *theirs*; 0 of 157 false invalidations; 189/189 FP16 checkpoint parity; 62 tests wiring not quality; 10× not achieved **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/validation.md#eval--hill-climb` |
| (meta) Finding new mappings & applications | Toolbox sweep: judgment-shaped component of a known method, substituted + falsified. **Hourly 1241:** groundedness-judge-bench native vs schema-guided; jev_playground 0 promotions; jsort scores are relative; jevbrain AUTO_ACT is not a Noul. **SIGNAL §94:** compile-time catalysts ≠ summaries; unofficial JA JGLUE *theirs*; NAR/multimodal *theirs*. **Hourly 1541:** difficulty + policy thresholds + JSONL trace; quarry evidence projection; one-dollar-tahoe TypeSafe Jev defense eval; llama-jev llama.cpp replica. **Hourly 1740:** Decision Graph Protocol frame→assess→commit; jegrep calibrated path+range Nouls; no embeddings/index/daemon; ~$0.01–0.03 typical; agent --json; Archer-arch fidelity; kev family OOD 0.76–0.77 vs Jev 0.86; replica honesty. **Hourly 1843:** cost-sensitive control flow overlay; variable-N option scoring as the trainable object; NAR local drop-in; typed vs chat judges on guardrailing. **Hourly 1943:** Jev IS the if-statement; retrieve by relevance not resemblance; memory leases ended by new evidence; JSON Schema → typed JSON via Jev **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/toolbox-mapping.md` |
| Named methods / operators / theorems | Substitution tiers: operand-judgments, preconditioned theorems, non-substitutable. **Hourly 1241:** ZHUBoer/ego-jev reserved `__none__`; jsort scores are relative; groundedness-judge-bench native vs schema-guided; jev_playground 0 promotions | `references/methods-catalog.md` |
| (meta) Where a judgment model sits relative to any construct | 11 positions + logical-operator rules + position×construct traversal as the application generator. **Hourly 1241 items 49–61** (`notes.md` §90). **SIGNAL §93 items 80–81** (`notes.md` §93): Decision ledger / memoization; GEPA alignment loop. **SIGNAL §94 items 82–84** (`notes.md` §94): Compile-time System One / questions-as-index; unofficial JA ModernBERT; NAR/multimodal/Router-OOD. **Hourly 1541 items 85–90** (`notes.md` §95). **Hourly 1639 items 91–92** (`notes.md` §96). **SIGNAL §97 item 93** (`notes.md` §97). **Hourly 1740 items 94–96** (`notes.md` §98). **Hourly 1843 items 97–101** (`notes.md` §99). **Hourly 1943 items 102–110** (`notes.md` §100) **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/composition-algebra.md` |
| Question mechanics & debugging | Instruction/criteria/state shape, budgets, diagnosis table, revision discipline. Missing `other` → confident wrong Choice (confidence gating cannot catch); lint the request (`wellposed` recipe; `tenbin` owns the skill). **Hourly 1144:** Treat `.feels()` 0.5 as a bool if / new language. **Hourly 1241:** Treat `completed` as success / Choice as the scale / 83% as quality / AUTO_ACT as a Noul. **SIGNAL §94:** Hard-gate enzyme `when asked` / treat catalysts as summaries; unofficial JA as TypeSafe; collapse LFM default into JA ModernBERT; treat enzyme hosted bootstrap as silent TypeSafe; Nemotron as calibrated Jev. **Hourly 1541:** Treat orchestrator “difficulty” as a live Score; invent one-dollar-tahoe ASR/FPR; treat llama-jev softmax as a Noul. **Hourly 1740:** Treat DGP as official TypeSafe / assessment p as a grant; collapse jegrep into jevgrep; treat kev OOD 0.76 as Jev; treat Archer as landed. **Hourly 1843:** Treat UNSURE as False; treat min_confidence as P(correct); treat gut/judge as new species; merge von Needle 52.6% with 93% n=78; paste Jev wins guardrailing. **Hourly 1943:** Treat Probably as hunch/gut/Judge/tidymodels; hard-gate feels 80%; treat jev-lint rename as a second product; paste 17/18 as Harbor; hard-gate 0 of 157; treat boolean @ 0.5 as a proof; claim 10×; treat softmax as a Noul; treat 62 tests as quality **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK | `references/question-design.md` |
| Heuristic search over a taxonomy | Parallel beam over Choice distributions | `references/mappings.md#5-hierarchy--bounded-heuristic-search` |
| Existing-system insertion / code-smell audit | Opportunity map, fit test, smallest boundary, policy centralization | `references/boundary-audit.md` |

Each card carries its boundary, counterexample, and acceptance test, plus
explicit rejections beside the mapping they tempt (MCTS-as-value-function:
experimental; bandits: rejected without observed rewards; 255-way tournament
brackets: rejected as default; rerank-huge-sets: budget-only; correlated
"independent" checks: rejected; listwise ranker *as* a fail-closed gate:
rejected; Noul *as* a proof / model-check / DST property: rejected;
TOCTOU-of-Noul as authorize: rejected; tautological spec + "looks good":
rejected). Promote Hypothesis cards only with an acceptance test that ran.

## Non-negotiable boundaries

- Score is an expectation over level indices, not a measurement in natural
  units. `[0,1,0]` and `[0.5,0,0.5]` both score 1.0 with different risk.
- Noul 0.5 is uncertainty about a predicate, never medium intensity.
- Parallel answers are not statistically independent: never multiply them
  into a joint probability.
- Choice probabilities are conditional on the offered set; absent candidates
  can never be chosen. No-match options (`other`) where coverage is open.
- Untrusted state text cannot authorize actions. Missing evidence is not
  evidence of absence. Validate operation+target pairs in code. Policy
  (checklist, ledger, law, two-person rule) is the code of a practice
  that has no repository.
- Never launder a Noul (or any judgment-class score) as a proof, a
  model-check, or a DST property. Soft check ≠ interlock. TOCTOU-shaped
  gates and tautological specs are harms, not placements. Exact work
  stays in code/policy; the model owns narrow judgment only. Full cards:
  `references/formal-methods.md`, `references/formal-semi-formal.md`,
  `references/mental-models.md`.
- Ranking scores order; decision scores authorize. Listwise / pairwise
  discriminative losses are translation-invariant: do not fail-closed on
  them. CLIP/SigLIP/GLiNER/GLiClass affinities are not automatically
  class-conditional P(permit). Full class card:
  `references/judgment-class.md`.
- Classification is not the product. The claim is a **placement**: typed,
  (when trained for it) calibrated judgments that policy can threshold —
  beside generation and beside exact work, not instead of them. Working
  regexes, ledgers, and recipes stay; trained classical classifiers still
  win on stable labeled taxonomies; open-ended writing stays generated.
  Full answer: `references/faq.md`.

## Decision-design card

```text
Domain (AI / SWE / business / knowledge work / life / org):
Desired behavior and non-judgment baseline:
Semantic judgment(s) and what each output means:
Pillar (EU / VOI / MCDA / SDT / search / safety / formal):
Hole (sieve / keep-drop / triage / rank / route / gate / perceive / abstain / gather):
Family (closed decision API / open head / open multimodal RLCD / encoder open-jev / constrained-AR surface / GLiNER locate / GLiClass categorize / listwise ranker / vision scorer):
Evidence/candidate source and known coverage gaps:
Deterministic policy, constraints, and action ownership:
Batchable vs genuinely dependent steps:
Failure/abstention behavior (fail-open vs fail-closed, matched to the family):
Smallest experiment that could reject this family, not just this vendor:
Eval path (jevals-shaped held-out and/or Harbor taskset; missing = incomplete):
Typed judgment provider (TypeSafe Jev default; other family only with self-eval):
Live references + versions (model, rubric, policy):
```

For open-ended requests propose three materially different *placements of
judgment*, recommend one. Cross-domain frames:
`references/mental-models.md`. Mixed-architecture extras:
`references/mixed-architecture.md`. Family-choice extras:
`references/judgment-class.md`. Formal / proof / DST:
`references/formal-methods.md`. Applied SWE placements (sieve, keep/drop,
env triage, moderation/ranking, skill routing):
`references/applied-mappings.md`. Hypothesis cards (VOI, SDT, Leveson,
search/control, spec pipeline, Alloy loop, RV sandwich, DST triage,
durable agents, assignment, situated density, paraphrase stability,
structural-prove ∩ remainder, effect-oriented state-machine loops)
stay labeled until an acceptance test
runs: `references/mappings.md` §6–§19. For concrete
requests skip the brainstorm and build.

## Evidence labels

- **Contract**: current documented behavior — refresh from live docs.
- **Empirical recipe**: worked on a stated dataset/model/version — retest.
- **Hypothesis**: plausible, unestablished — label and test before relying.