log-analyzer · diff

v1.0.0 to v1.1.0

58 added, 9 removed. Audit A to A.

---
name: log-analyzer
- description: "Parse agent log files to identify error patterns, rate limit hits, timeout clusters, tool failures, and component-level error counts. Produces a structured anomaly report. Cron-compatible — silent if no issues, alert digest if anomalies found."
+ description: "Parse agent log files to identify error patterns, rate limit hits, timeout clusters, tool failures, and component-level error counts. Produces a structured anomaly report. Cron-compatible — silent if no issues, alert digest if anomalies found. Also computes per-tool failure rates from a Hermes profile state.db (scripts/state_failures.py)."
license: MIT
metadata:
- version: 1.0.0
+ version: 1.1.0
author: moonlight-lupin
platforms: [linux, macos, windows]
tags: [logs, analysis, errors, patterns, anomalies, monitoring, debugging, cron]
related_skills: [scheduled-summary, skill-maintainer]
---
# Log Analyzer
## Overview
Log viewers filter lines. This skill finds patterns.
Use it when an agent runtime, tool process, gateway, scheduler, or other service
has produced enough log output that individual `grep` hits no longer explain the
system behavior. The analyzer parses standard log lines, normalizes repeated
messages, groups errors by component and tool, and produces a structured anomaly
report that is suitable for debugging sessions or cron digests.
The script is intentionally lightweight and portable: it uses only Python's
standard library and works on any text log with timestamp, level, optional
component, and message fields. It also has a best-effort fallback for
unstructured logs.
## Quick Start
```bash
cd agent-ops/log-analyzer
python scripts/analyze_logs.py scan --log-file agent.log --since 24h
```
+ For per-tool failure rates from a Hermes profile's session DB (structured
+ exit_code signals, NOT regex-over-content):
+
+ ```bash
+ python scripts/state_failures.py # last 7 days, dashboard
+ python scripts/state_failures.py --days 30 --json
+ python scripts/state_failures.py --quiet # cron: silent when healthy
+ ```
+
+ See `references/state-failure-monitor.md` for provenance (adapted concept
+ from hermes-dojo), the spike evidence for why structured signals matter,
+ and interpretation notes.
+
To write JSON for later Markdown rendering:
```bash
python scripts/analyze_logs.py scan --log-file agent.log --since 24h --output scan.json
python scripts/analyze_logs.py report --scan scan.json --output report.md
```
For cron-compatible anomaly detection:
```bash
python scripts/analyze_logs.py scan --log-file agent.log --since 24h --quiet
```
`--quiet` exits 0 and prints nothing when no anomalies are found. If anomalies
exist, it prints the JSON report so the scheduler can deliver the digest.
## What It Detects
1. **Error clusters** — the same normalized error message repeated 3+ times
within the selected time window. URLs, IP addresses, and numbers are replaced
with placeholders before grouping so repeated failures with changing IDs still
cluster.
2. **Rate limit hits** — HTTP `429`, `rate limit`, `rate_limit`, `too many
requests`, and `quota exceeded` patterns. The analyzer groups them by
provider when it can detect a provider name.
3. **Timeout patterns** — `timeout`, `timed out`, `deadline exceeded`, and
`connection timeout`. Results are grouped by detected tool name and include
example URLs where present.
4. **Tool failures** — error lines grouped by tool name extracted from patterns
such as `tool: terminal`, `tool_call: web_search`, `tool terminal failed`, or
lines emitted by a `tools` component.
5. **Session crashes** — fatal errors, unhandled exceptions, stack traces,
`Traceback`, `Exception`, and `segfault` markers. Multiline stack traces are
grouped as one crash entry with nearby context.
6. **Component breakdown** — error and warning counts by component such as
`gateway`, `agent`, `tools`, `cron`, or `unknown`.
7. **Error timeline** — error counts bucketed by hour to reveal spikes and
regressions after deploys or scheduled jobs.
See `references/anomaly-types.md` for detection criteria and interpretation.
## Log Format Support
The parser handles standard log lines shaped like:
```text
2026-07-06 12:30:45 ERROR [gateway] Connection refused
2026-07-06T12:30:45Z ERROR gateway: Connection refused
12:30:45 ERROR Connection refused
```
It recognizes `ERROR`, `WARN`, `WARNING`, `INFO`, `DEBUG`, `FATAL`, and
`CRITICAL` levels. Components may appear in square brackets after the level or as
`component:` after the level. Time-only lines are anchored to the current day (or
to the supplied default date when called as a library). Unstructured lines fall
back to best-effort line-by-line scanning, so obvious `ERROR`/`WARN` strings are
still counted even when the timestamp cannot be parsed.
See `references/log-formats.md` for examples and guidance on adding custom
patterns.
## CLI Commands
### `scan` — analyze a log file for patterns
```bash
python scripts/analyze_logs.py scan --log-file LOGFILE [--since TIME] [--output report.json] [--quiet]
```
Options:
- `--log-file LOGFILE` — required path to the log file.
- `--since TIME` — optional time window: minutes/hours/days/weeks, e.g. `30m`,
`1h`, `24h`, `7d`, `2w`; default is all lines.
- `--output report.json` — write JSON to a file instead of stdout.
- `--quiet` — cron mode: suppress output when no anomalies are found.
### `report` — render Markdown from scan JSON
```bash
python scripts/analyze_logs.py report --scan scan.json [--output report.md]
```
The report contains overview counts, one section per anomaly type, a component
breakdown table, and an hourly error timeline.
### `tail` — smart tail for recent lines
```bash
python scripts/analyze_logs.py tail --log-file LOGFILE [--lines N] [--since TIME]
```
The smart tail prints plain text with markers suitable for chat delivery:
- normal lines: no marker
- warnings: `⚠️`
- errors: `❌`
- repeated recent errors: `🔥` when the same normalized error is seen 3+ times
## Output Format
`scan` emits JSON with these top-level fields:
- `log_file`, `lines_analyzed`, `time_window`
- `anomalies.error_clusters`, `anomalies.rate_limits`, `anomalies.timeouts`,
`anomalies.tool_failures`, `anomalies.crashes`
- `component_breakdown`
- `error_timeline`
- `total_errors`, `total_warnings`, `has_anomalies`
`report` converts that JSON into Markdown:
```markdown
# Log Analysis Report
## Overview
- Log file: agent.log
- Lines analyzed: 1542
- Time window: 24h
- Total errors: 12 | warnings: 20
- Anomalies detected: 5 types
```
## Scheduled Summary Integration
For a scheduled digest, run `scan` in quiet mode and include the output only when
it is non-empty:
```bash
python scripts/analyze_logs.py scan --log-file /var/log/agent.log --since 24h --quiet --output /tmp/log-scan.json
if [ -s /tmp/log-scan.json ]; then
python scripts/analyze_logs.py report --scan /tmp/log-scan.json
fi
```
A `scheduled-summary` job can append the Markdown output under a "Log anomalies"
heading. Keep the analysis window aligned with the summary window (for example,
24 hours for a daily digest) so counts do not overlap or disappear.
+ For a state.db failure-rate digest (weekly is a sensible cadence given the
+ volume of sessions):
+
+ ```bash
+ python scripts/state_failures.py --quiet --days 7
+ ```
+
+ This prints nothing (exit 0) when the window has zero failures, so it can be
+ used as a `no_agent` cron job that only pings when something is wrong. See
+ `references/state-failure-monitor.md` for what the categories mean (timeout
+ is the most actionable on real data).
+
## Common Pitfalls
1. **Log rotation breaks time windows.** If yesterday's file was rotated out, a
`--since 24h` scan over only the current file may miss early-window failures.
Point the scheduler at the active file plus rotated file, or concatenate the
relevant files before scanning.
2. **Multiline stack traces need the first line.** The parser groups indented
stack-trace continuation lines under the preceding parsed log line. If a log
collector strips the first `Traceback` or `ERROR` line, the remaining stack
frames become unstructured context.
3. **Expected errors can be false positives.** Retries, probing, and health
checks may intentionally emit warnings or connection failures. Treat clusters
as "investigate" signals, not automatic incidents.
4. **First run on a large historical log can overwhelm output.** Start with
`--since 24h`, inspect the report, then widen the window if needed.
5. **Changing message formats can split clusters.** If an application changes an
error string during a deploy, pre- and post-deploy failures may appear as two
clusters even when the root cause is the same.
6. **Time-only logs depend on the scan date.** `12:30:45 ERROR ...` lines do not
contain a date. For historical files, prefer full timestamps.
- 7. **The error timeline buckets by hour-of-day, not date.** On a multi-day
+ 7. **Comma-millisecond Python-logging lines silently defeat `--since`.** Lines
+ shaped `2026-07-06 12:30:45,123 INFO module: ...` (Hermes agent/gateway/errors
+ logs) are NOT matched by the built-in timestamp regexes. They fall to the
+ unstructured fallback with `timestamp=null`, so `--since` cannot filter them
+ and the **whole unrotated file is scanned**. Symptoms: error/traceback counts
+ wildly exceed the true windowed count, every cluster reports
+ "First: unknown / Last: unknown", the error timeline says "No timestamped
+ errors detected", and crashes are inflated (each stack-trace continuation
+ line counts as a separate crash). **Fix:** for Hermes profile logs use the
+ windowed counter `scripts/hermes_log_window.py` (handles the comma-ms format,
+ de-duplicates multiline tracebacks by counting `Traceback (most recent call
+ last)` first-lines only) and see `references/hermes-profile-audit.md` for
+ pulling skill-usage/tool-call counts from the profile's `state.db`.
+ 8. **Skill-usage counts are not in agent.log.** `tool skill_view completed
+ (0.05s, 13890 chars)` lines do not carry the skill-name argument. Query the
+ profile's `state.db` `messages` table — the `tool_calls` JSON column holds
+ the full arguments. `messages.timestamp` is a Unix epoch (REAL), not ISO.
+ See `references/hermes-profile-audit.md` for a ready-to-run snippet.
+ 9. **The error timeline buckets by hour-of-day, not date.** On a multi-day
window (`--since 7d`), errors from 08:00 on different days merge into one
`08:00` bar. Use the clusters (which carry full timestamps) for multi-day
forensics; treat the timeline as a time-of-day profile.
- 8. **Crash detection is substring-based.** The crash regex matches words like
- `fatal` or `Exception` anywhere in a line, so mentions inside INFO lines
- (e.g. "retry succeeded after TimeoutException", "non-fatal warning") count
- as crash signals and can set `has_anomalies`. Treat crash counts as leads
- to eyeball, and tune the regex if your logs legitimately chat about
- exceptions at INFO level.
+ 10. **Crash detection is substring-based.** The crash regex matches words like
+ `fatal` or `Exception` anywhere in a line, so mentions inside INFO lines
+ (e.g. "retry succeeded after TimeoutException", "non-fatal warning") count
+ as crash signals and can set `has_anomalies`. Treat crash counts as leads
+ to eyeball, and tune the regex if your logs legitimately chat about
+ exceptions at INFO level.
## What This Skill Is NOT
- Not a log viewer: use `tail`, `less`, or a log UI when you need raw line
inspection.
- Not a log shipper: it does not forward logs to storage or observability
systems.
- Not a SIEM: it does not correlate identities, networks, or security events.
- Not real-time monitoring: `tail` is a recent-line analyzer, not a daemon or
alerting service.
- Not a root-cause oracle: it highlights patterns so an agent or operator can
investigate faster.
## Verification Checklist
- [ ] Run `python scripts/analyze_logs.py --help` and confirm subcommands load.
- [ ] Scan a synthetic log with repeated errors and confirm `has_anomalies` is
true.
- [ ] Render Markdown from the scan JSON and confirm all anomaly sections appear.
- [ ] Run `python -m pytest tests/test_analyze_logs.py -v` from this skill
directory.
+ - [ ] Run `python -m pytest tests/test_state_failures.py -v` from this skill
+ directory (state.db failure-rate monitor; the spike regression test —
+ content merely mentioning error words must not count as failure — lives
+ here).
- [ ] For cron use, test a no-anomaly log with `--quiet` and confirm stdout is
empty with exit code 0.
+ - [ ] For state.db cron use, `python scripts/state_failures.py --quiet` against
+ a live profile DB exits 0 and prints nothing when the window is clean.