AGENTS.md · git:20260906.5f1fa6f · 2026-09-06 · sha256 8da4cf819934e7a7
AGENTS.md git:20260906.5f1fa6fA
Immutable. This exact content is served forever at /api/v1/blob/8da4cf819934e7a7.
# AGENTS.md — working on oh-my-design with Codex (GPT-5.6) This file governs how the Codex / GPT-5.6 path contributes to this repository. `CLAUDE.md` covers the Claude Code path; both share the repository conventions at the bottom. ## Model ownership: the user chooses the model and any invocation override The concrete model selected for the host session is user-owned configuration. OMD never replaces it for a child agent unless the user supplies an explicit role override on that Codex host invocation. This applies equally to Codex and Claude Code, to every pipeline role, and to ad-hoc workers spawned during the run. - **Codex:** omit `model` from every `spawn_agent` call. Official brokered roles also omit a model unless the user supplied that role's host-only `--omd-role-model` option on this invocation. - **Claude Code:** agent metadata declares `model: inherit`; never request Opus, Sonnet, Haiku, or any concrete model in a spawn. - **Both hosts:** source-owned defaults set only the role's reasoning/effort tier. Judgment-heavy roles use `high`; the production hand uses `medium`. A user may override a Codex role's effort for one host invocation with `--omd-role-effort`. - A recommendation, benchmark, role name, or belief that another model would perform better is never authority to override the user's selection. Without an explicit host override, a Luna session keeps Luna for every OMD child and a Sol session keeps Sol. The pipeline's agent source files carry only the effort tier. The Codex adapter normally emits `model_reasoning_effort` and no `model`; the Claude adapter emits `model: inherit` plus `effort`. Codex role overrides are valid only when the user places `--omd-role-model role=model` or `--omd-role-effort role=low|medium|high` on the outer `omd-codex exec` command. The host strips those options before coordinator launch, binds them to that invocation, and applies them inside the broker. A coordinator, role task, or installed agent profile must not invent an override. ## Adaptive design and evidence ownership - Host wrappers preserve each role's declared publication method. Direct writes apply only to explicitly owned, directly editable paths and never replace a required CLI publisher. Complete current evaluator lineage belongs in the public input skeleton and the source-free judgment packet; transport repairs do not authorize reconstructed judgments or weaker evidence checks. - Acquisition `requiredState` describes an inspectable reference component state, not the destination's promised behavior. Keep destination outcomes in requirements and falsifiers. Framer repairs a misframed plan; Scout never relabels an observation to pass exact state binding. - Sparse reference pages are not automatically blocked pages. A successful HTTP response with a measured, visible, nonempty scoped component may resolve the short-body heuristic; document-root selectors, empty/hidden content, HTTP 403/server errors, and challenge titles do not gain an exemption. - The outcome, risk, uncertainty, and available evidence select the route. Optional stages and methods require either selection or a written skip; the full capability catalog is not a universal sequence. - Concept exploration records its candidate count and evidence-based reason before generation. Distinct concepts need different content-to-form relationships and macro composition, not different brand colours. Anchor and background counts follow content. An unshipped image draft can still be useful decision material; asset shipping restrictions do not by themselves justify skipping it. Select feasible concepts for rendered task fit and craft; use cost only for an explicit budget constraint or a tie between equivalent candidates. `core/theory/imagegen.md` owns the procedure. - `new-product` and `new-marketing` are distinct reference-discovery needs. Greenfield product work can require a task-flow benchmark; a marketing launch does not impersonate a product workflow. - A prompt-only greenfield marketing brief that demands a product difference but supplies no verified capability or mechanism stops before art direction and production. Category references identify missing facts; they do not become destination capabilities. - Benchmark product evaluation begins from the current task outcome, evidence claims, frame-owned selector-free entry contract, and sanitized task-flow projection. `omd lifecycle plan` derives fixed evaluator selectors and exact assertions; callers cannot substitute a different manifest. - Entry contracts support both prerequisite → gated-action flows and prerequisite → observable- consequence flows. A next action is required only when the task actually has one. The work object must begin at the entry viewport with its purpose and representative anchor visible; a legitimate long mobile work object may continue below the fold. - Real browser actions and fixed desktop/mobile captures prove behavior and entry fitness. Visual quality remains an isolated proxy review unless the human-calibration protocol has real blind designer and target-user ratings. - Locale-grounded design keeps conversation language, surface locale, explicit market region, audience/task, domain, surface, desired fit, and brand invariants separate in the sole writable `.omd/locale-design-context.json`. A locale or likely script may select mechanics; it never infers a market, register, or country aesthetic. Missing market/audience authority yields one question. - `mechanics-only` proves real target-language type without a cultural-fit claim. `research` binds current standard, global-equivalent or unavailable, native first-party, and counterexample source receipts into a content-addressed profile. Composer, Hand, and Eye receive only the source-free projection. `omd locale profile-check` re-hashes current type/source records and re-fetches remote sources; a changed context, proof, receipt, source, or projection fails closed. - A global equivalent must serve the same named task/category; a same-owner homepage is insufficient. Recorded unavailability limits confidence and never contributes to `supported`/`shared` convergence. - `supported` and `shared` profile mechanisms may transfer. `contested` needs a downstream decision; `unknown` adds no design rule. Profile compliance is evidence-grounded adaptation, not native cultural correctness; only blind ratings from the named target audience may support that claim. ## Repository conventions - **Source of truth is `src/`.** `src/agents/*.agent.yaml` and `src/skills/omd-*/SKILL.md` are the prompt originals. Root `agents/`, `skills/`, and `dist/` are generated by `npm run build` — never edit them directly. - **Edited directly:** `core/`, `bin/`, `adapters/`, `test/`, `evals/`, `scripts/`, `README.md`, `README.ko.md`, `.github/`, and the theory/recipe packs under `core/`. - **`.omd/` is the design record**, committed with the repo. `.omc/` is local session state and is gitignored. - **New linter rules stay narrow.** Every pattern gets a positive AND a negative test; a tell that cannot be made safe goes to the prompt layer with the reason recorded in-file. Rules warn, never error — a deliberate choice can overrule one with a written reason. - **Commits use conventional prefixes** (`feat:`, `fix:`, `docs:`, `chore:`). Do not add AI attribution footers to commit messages. - **Branch → PR → squash-merge.** Every change lands on a feature branch and merges through a PR; `main` is protected (one approval; the admin, 3x-haust, may bypass). - **Merge continuously, release on request only.** Feature branches merge to `main` as they finish — do not cut a version for every feature. A release happens **only when the user explicitly asks** for one. Work accumulates on `main` between releases. - **Releasing (when asked):** bump all three manifests with `scripts/bump.ts`, merge the bump to `main`, and the release workflow tags and publishes structured notes automatically. - **Definition of done for any change:** `npm test` 0 fail, `npx tsc --noEmit` clean, `npm run build` succeeds.