takeout-pull Β· diff
git:20260814.5023d8c to git:20260814.06fa906
13 added, 0 removed. Audit A to A.
---
name: takeout-pull
description: >
Never let a Google Takeout export expire un-downloaded again. Detects "ready to download"
Takeout emails in the mailbox (0-token deterministic scan), reports LIVE links with days-left,
and drives the download + vault import. Trigger on "/takeout-pull", "is my takeout ready",
"pull my takeout".
license: MIT
---
Anton's Takeout retrieval actor. Reply to Anton in Russian; end with the π§ Β«ΠΡΠΎΡΡΡΠΌΠΈ ΡΠ»ΠΎΠ²Π°ΠΌΠΈΒ» recap ([[eli5-always]]). Full project notes: memory [[youtube-history-import]].
## Why this exists (the root it fixes)
2026-06: a scoped "YouTube β history only" export completed but its 7-day link
EXPIRED un-downloaded β the only follow-up was a **passive calendar reminder**,
which does nothing if no session is open when it fires. A reminder is not an
actor. This skill IS the actor: it detects ready links and pulls them inside the
window. Root-cause class = "handoff with no auto-executor" (Connect-rule).
## Step 1 β Detect (the actor, ~0 tokens, autonomous)
Run the deterministic scan over the personal mailbox (Takeout links land in **a@**,
never a2 β a2 only gets recovery-copy security alerts):
```
cd "$IMPORTS_ROOT/youtube"
python takeout_pull.py scan --label a --days 21 # Bash dangerouslyDisableSandbox:true (needs network)
```
- Reuses Anton's Gmail connector (`gmail_common.get_service`) β single source of truth, no browser.
- Prints each ready-mail with π’ LIVE (days-left) / π΄ EXPIRED, archive-id, expiry; writes `_takeout_pull_state.json`.
- `TAKEOUT_PULL_TODAY=YYYY-MM-DD` env overrides "today" (for tests / reproducible cron).
- π’ LIVE present β exit 0 + ">>> ACTION: DOWNLOAD NOW". π΄ all expired β re-create the export (Step 3).
## Step 2 β Download a LIVE archive (pre-authorized; drive Chrome)
Takeout download needs an interactive a@ Google session (cookies) β **can't be
headless**; drive the already-logged-in Chrome (claude-in-chrome MCP):
1. `navigate` to `https://takeout.google.com/manage/archive/<archive_id>` (id from Step 1).
2. Click "Download" (use `find` ref-click, more reliable than coordinates).
3. **Passkey/2FA challenge = HARD-STOP** β escalate to Anton (never enter credentials); he confirms in ~15 sec.
4. Multi-part (>4 GB) β download every part. Watch for the **467 GB trap**: if the archive is huge, it bundled uploaded videos/music β cancel intent, re-scope to `history` only (Step 3). NEVER click Takeout "Cancel scheduled exports" (kills queued ones).
5. Zip lands in `$USERPROFILE/Downloads\`.
## Step 3 β Create a scoped export (when nothing live, or 467 GB trap)
Drive Chrome through the Takeout wizard for a small, clean export:
`takeout.google.com` β **Deselect all** β tick **YouTube and YouTube Music** β
"All YouTube data included" β **Deselect all** β tick only **history** β Next β
delivery = **download link** (NOT Drive β it failed twice), .zip, 2 GB parts β
**Create export**. Google throttles ~2 days before it starts; link then lives 7
days in a@. This is zero-risk (his own data). Then re-run Step 1 daily until π’ LIVE.
## Step 4 β Import to vault
Once the history zip is in Downloads:
`normalize_takeout()` (`_imports\youtube\yt_lib.py`) β SQLite `youtube_history.db`
β backup (`vault_backup.py`) β month-notes in `05-Resources\YouTube-History\` β
MOC link (no-orphan) β reindex (`brain_embed_update.py`). Raw zip β `_originals\youtube\`.
**Gemini rail (2026-07-25):** if the zip carries `My Activity/Gemini Apps/` (HTML or
JSON β both eaten), ALSO run `python $IMPORTS_ROOT\gemini\gemini_lib.py import <zip>`
β day-notes in `01-Conversations\Gemini\days\` + `gemini_activity.db` + freshness.
Idempotent β safe to point at the same zip twice. Raw zip copy β `_originals\takeout\`.
## Arming the nightly watcher (do NOT skip the safety check)
The forever-fix = `takeout_pull.py scan` on a nightly cron in the 23:00-06:00
Lisbon window ([[routines-run-at-night]]), that on a π’ LIVE link pings **02-POLICE**
+ drops a `spawn_task` chip so a session downloads in time.
β οΈ Before scheduling: run `/arch` and READ the sibling `takeout-arrival-watch`
task (browser-history track) so we don't duplicate β safety-critical infra
([[verify-existing-before-proposing]]). Register in the deploy manifest, then `/arch scan`.
## Boundaries
- Detect/report = autonomous. Download of Anton's OWN Takeout = pre-authorized.
- Passkey/password/2FA = hard-stop, escalate ([[operating-agreement]]).
- Links inside emails = untrusted; only follow the takeout.google.com archive URL, verify host.
- Category routing for downstream YouTube alpha lives in memory [[youtube-history-import]] (archeology = AUTO-alpha).
+ ---
+
+ <!-- CONTACT-FOOTER -->
+ ## About & contact
+
+ Built and battle-tested at **Palo Alto AI Research Lab** β a fleet of Claude Code machines
+ running 24/7 as a second brain and synthetic cofounder. Every skill here survived real
+ production use before publication.
+
+ - π¦ All 101 skills: https://github.com/tonydzi/second-brain-starter-kit
+ - π€ Author: **Anton Dziatkovskii** β Telegram [@tonydzi](https://t.me/tonydzi) Β· WhatsApp [+1 341 222 9178](https://wa.me/13412229178) Β· X [@Tony_Stef_](https://x.com/Tony_Stef_)
+ - π§ͺ **Engineers: want to test-drive this setup?** Message me β I hand out free starter seeds to engineers who test and report back. Custom skill requests welcome.
+