AGENTS.md ยท diff

git:20260513.58ebf0e to git:20260515.025f516

28 added, 1 removed. Audit B to B.

# Agent Instructions
This file is repo-local guidance for Codex, Claude Code, and other coding
agents working on Eval Engineer. Keep it short and update it whenever the repo
workflow, source-of-truth files, or recurring operating rules change.
## Project Intent
Eval Engineer is a general Galileo evidence workflow for improving AI agents,
RAG apps, and future AI systems. Do not shape the product around the first
support-agent fixture. The support-agent cases are validation fixtures, not the
skill's scope.
The near-term product goal is to reduce time-to-RCA for agent builders and
owners while increasing Galileo discoverability for non-developer personas.
Prefer RCA workflows that query log streams, inspect traces/sessions/spans,
identify failure patterns, compare behavior over time, and return grounded
answers with links or stable IDs back to Galileo data.
The north-star loop is:
1. run the AI app
2. log traces and metrics to Galileo
3. fetch compact evidence
4. diagnose failure
5. propose a bounded fix
6. verify with local and Galileo evidence
7. keep only changes that improve measured behavior
## Read First
- `docs/plan.md`: product direction and architecture.
- `docs/tasks.md`: current task checklist and Linear issue mapping.
- `docs/progress.md`: latest work completed and next move.
- `.galileo/learnings.md`: repo-specific durable learnings.
- `blogs/`: product thinking; useful for design intent, but do not treat as
runtime instructions.
## Skill Rules
- Canonical skill source: `skills/galileo-eval-engineer/`.
- Codex install link: `.agents/skills/galileo-eval-engineer`.
- Claude install link: `.claude/skills/galileo-eval-engineer`.
- Keep `SKILL.md` general across agents, RAG, workflows, metrics, and providers.
- Keep RCA outputs grounded in trace, span, session, metric, dataset, and
experiment evidence.
+ - For reference fixtures, give each case a risk profile, quality dimensions,
+ and case-specific Galileo metric profile. Do not rely on one global metric
+ list to prove safety, quality, performance, and cost across all cases.
+ - Use `skills/galileo-eval-engineer/references/metric-profile-checklist.md`
+ and `skills/galileo-eval-engineer/assets/metric-profile-template.md` before
+ optimizing cost or adding broad fixture coverage.
+ - When logging test-suite cases to Galileo, include the full expected-output
+ contract in `dataset_output`: expected decision, required/forbidden citations,
+ required tools, answer constraints, abstention/permission requirements, risk
+ profile, quality dimensions, and intended metrics.
- Use `skills/galileo-eval-engineer/references/rca-recipe.md` for generalized
diagnose-fix-verify work and update it when a reusable Galileo RCA pattern is
discovered.
- Keep detailed Galileo mechanics in `skills/galileo-eval-engineer/references/`.
- Keep deterministic helpers in `skills/galileo-eval-engineer/scripts/`.
- Do not hardcode `TC-1`, the Nexus support agent, Brazil, one model, or one
metric into the general skill.
## Working Set
- `.galileo/config.yml`: agent type, metrics, editable files, verification
commands.
- `.galileo/current/`: current evidence and working artifacts.
- `.galileo/sessions/`: historical evidence.
- `.galileo/eval-dataset/`: candidate, accepted, and rejected eval cases.
- `.galileo/learnings.md`: durable patterns discovered while working.
Read `.galileo/current/` by default. Do not scan historical raw sessions unless
the user asks for history or comparison.
## Verification
After skill changes, run:
```bash
PYTHONPYCACHEPREFIX=/private/tmp/eval-engineer-pycache python3 -m unittest tests.skills.test_galileo_eval_engineer_skill
python3 /Users/pratik/.codex/skills/.system/skill-creator/scripts/quick_validate.py skills/galileo-eval-engineer
```
For the generic packet summarizer:
```bash
python3 skills/galileo-eval-engineer/scripts/summarize_debug_packet.py tests/skills/fixtures/generic-rag-debug-packet.json
```
+ For tokenomics before/after packet comparison:
+
+ ```bash
+ python3 skills/galileo-eval-engineer/scripts/compare_tokenomics_packets.py .galileo/current/debug-packet.json .galileo/current/verification-debug-packet.json --quality-metrics average_tool_selection_quality,tool_error_rate,average_completeness_gpt,average_groundedness --lower-is-better-quality-metrics tool_error_rate
+ ```
+
+ For tokenomics work, compare cost, latency, and token movement against Galileo
+ quality metrics. Do not keep a cost reduction on local scoring alone, and do
+ not treat lower traffic volume as a per-trace efficiency improvement. Check
+ segment metrics when available before accepting an aggregate cost win. For RAG
+ retrieval pruning, include at least one hard multi-source or multi-hop case so
+ top-1 retrieval shortcuts cannot pass on easy single-document questions alone.
+ For agentic workflows, also compare agent steps, planner spans, rerank passes,
+ self-check spans, and tool calls so an adaptive optimization shows which part
+ of the loop got cheaper.
+
Use live model or Galileo calls only when the task explicitly requires runtime
verification. Do not print secret values.
## Linear Hygiene
Keep Linear up to date manually after response completion when task status
changes. There are no repo hooks for this.
- Create Linear issues for new meaningful work.
- Move completed Linear issues to Done.
- Move obsolete issues to Canceled or Backlog with an explanatory comment.
- Keep `docs/tasks.md` issue IDs aligned with Linear.
- Mention the relevant Linear IDs in `docs/progress.md`.
Current active planning issues:
- `GAL-87`: answer-quality check for policy explanation correctness.
- - `GAL-82`: add first RAG reference implementation.
+ - `GAL-93`: Eval Engineer launch blog.
+ - `GAL-94`: Eval Engineer launch graphics.
## Secrets And Generated Files
- Never print `.env` values. It is okay to list variable names.
- Keep generated runs, raw traces, temporary logs, and `.omc/` out of git.
- Preserve user changes. Do not revert unrelated edits.
## Maintenance Rule
If you discover a new durable workflow rule, repeated mistake, packaging
decision, or verification command, update this `AGENTS.md` in the same change.