hive.error-recovery · diff
git:20260318.0cef0e6 to git:20260416.e719523
18 added, 7 removed. Audit A to A.
---
name: hive.error-recovery
- description: Follow a structured recovery protocol when tool calls fail instead of blindly retrying or giving up.
+ description: Follow a structured recovery decision tree when tool calls fail instead of blindly retrying or giving up.
metadata:
author: hive
type: default-skill
---
## Operational Protocol: Error Recovery
When a tool call fails:
- 1. Diagnose — record error in notes, classify as transient or structural
- 2. Decide — transient: retry once. Structural fixable: fix and retry.
- Structural unfixable: record as failed, move to next item.
- Blocking all progress: record escalation note.
- 3. Adapt — if same tool failed {{max_retries_per_tool}}+ times, stop using it and find alternative.
- Update plan in notes. Never silently drop the failed item.
+ 1. **Diagnose** — classify the failure as *transient* (network blip, rate limit, timeout) or *structural* (wrong selector, missing auth, invalid schema, permission denied).
+
+ 2. **Decide:**
+ - Transient → retry once.
+ - Structural + fixable → fix the input and retry.
+ - Structural + unfixable → record the failure and move to the next item.
+ - Blocking all progress → escalate.
+
+ 3. **Adapt** — if the same tool has failed {{max_retries_per_tool}}+ times in a row, stop using it and find an alternative approach.
+
+ **Never silently drop a failed item.** If the item is a task in the colony queue, write the failure to the DB instead of an in-memory buffer:
+
+ ```bash
+ sqlite3 "$DB_PATH" "UPDATE tasks SET status='failed', last_error='<one-sentence reason>', completed_at=datetime('now'), updated_at=datetime('now') WHERE id='<task-id>' AND worker_id='<your-worker-id>';"
+ ```
+
+ The `tasks.retry_count` column and the stale-claim reclaimer handle auto-retry for crashes; your job is the within-run decision tree above. See `hive.colony-progress-tracker` for the full queue protocol.