---
name: spot-audio-assembly
description: >
  Assembles the SOUND of a generated-video piece: voice-over takes (one voice and model per role), dubs for
  refused lines in the character's cloned voice, off-screen lines, sfx with the syllable map confirmed before the
  cut, music cues cut to markers with ducks, room tone, captions, the stem, and the read of the delivered audio.
  Use when a line must be voiced or re-voiced, a line was refused, sfx or music must be sourced or cut, captions
  must be built, a take's audio carries an artifact, or the delivered audio must be verified. Triggers — "generate
  the VO", "same voice as before", "dub the line", "the off-screen voice", "the sting", "the syllable map", "the
  music is too loud", "duck the music", "room tone", "captions", "is the vo on time", "musical artifact at the
  start". Not for placing lines on the timeline or the beat gate — use video-edit-edl. Not for the master render
  or the delivery encode — use video-finish-qc. Not for LUFS mastering of a finished mix — use mastering-audio.
allowed-tools: Read, Glob, Grep, Write, Edit, Bash(python3*), Bash(ffmpeg*), Bash(ffprobe*), Bash(ls*), Bash(curl*)
---

# Spot Audio Assembly

Sound is built from the picture and paid for by the character, the minute and the generation — so every
voice, sound and cue arrives through a **cost line and the operator's GO**, lands under a stable name
that is never regenerated by accident, and is placed by a **reference** the edit holds
(`video-edit-edl`). The operator's ear is the pick; the instruments locate, they do not decide.

**What varies.** Voices, models and the music route are vendor choices; the loudness target and the true-peak
ceiling are the client reference's or the platform's (the EDL carries both); the caption style is the campaign's; the
music licence follows the client's terms; the gap floor is a lead from one job. The cost line, the ordinal pick and the
read of the delivered file do not move — `video-production/references/WHAT-VARIES.md` § Audio and voice.

## 1. Cast the voices — one voice, one model, per role

Before the first take: the narrator's voice id (the operator's pick), the model that renders it for the
campaign's whole life, and a distinct voice for every other speaking role.

- **One model per voice.** `eleven_v3` renders the same voice id with a different timbre; identity is
  audited by ear ("is it the same voice?") — same id, same model, same settings, a `seed` when pinned.
- **An off-screen line is never the narrator** (the client hears the narrator asking a question that
  belongs to the character behind the door): a shared-library voice picked from a reel against the client's brief,
  rendered with the **door treatment** (`gen_vo.py --door`).
- **A line the venue refused** (profanity, an R-rated line) is **dubbed in the character's own
  cloned voice**: an instant clone from that character's take lines (every sample ≥ 4.6 s), the shot
  prompted with a similar-mouth word, the real line rendered by the clone. A real person's
  voice as a clone = likeness exposure, stated ONCE. The clone id is recorded and reused.

**Done when:** every speaking role has a voice id + model on record, the O.S. role is not the narrator,
and any clone's likeness note has been given.

## 2. Render takes through the cost line

```bash
python3 ~/.claude/skills/spot-audio-assembly/scripts/gen_vo.py --root <project> --jobs audio/vo/jobs-<SPOT>-<line>.json --dry-run
#   → characters per job and the total → the cost line to the operator → GO →
python3 ~/.claude/skills/spot-audio-assembly/scripts/gen_vo.py --root <project> --jobs … --preview audio/vo/PREVIEW-<SPOT>-<line>.mp3 [--door]
```

🔴 The cost line and the GO precede **every** TTS, STT, sound-generation and clone call, in
`video-gen-cost-gate`'s form, however small the amount. The script prints the
plan's character count before and after so the billed delta is on the record; an existing output is
skipped by name; a new take is a new job name.

- Three takes per line (stability / style / seed spread) into a **named reel with an index**; the
  operator picks **by ordinal from the named file**. "not bad" is not a pick.
- Steer the reading without speaking it: `prev`/`next` context (slang read as another language without
  it), phonetic spelling ("lotta", never "lot of"), a spelling spread across the takes.
- The full VO script is the `text` table in the EDL, quoted in every review ask.

Venue facts, auth traps, the dub and O.S. procedures: [`references/VO-PIPELINE.md`](references/VO-PIPELINE.md).

**Done when:** the pick is named by ordinal and file, its job is on disk under its final name, and the
billed delta is in the receipt.

## 3. Source the sound — the syllable map before the cut

The ladder: the take's own sound (cut at the frame, placed as an sfx — `video-edit-edl`), then a library
sound with its licence line read, then generated sound (billed → cost line). RECORDED sound — documentary
speech, an interview, an event's floor — carries a **microphone-position column** on every window (on the
subject · near the subject · far, on the floor): a window used under OTHER picture is subject-near, or it is
shown with its own frames, and a sound with a visible cause (a bystander's call, a ringman's yelp) is shown
with its picture or not used — a floor recording laid under other shots "sounds weird without context", and
the operator sent the build back (2026-09-15). A reel pick is confirmed as
a **syllable map before cutting** — which onsets, in which order, the bitten-off one excluded:

```bash
python3 ~/.claude/skills/spot-audio-assembly/scripts/sfx_onsets.py audio/sfx/<reel>.wav          # ordinals + times
python3 ~/.claude/skills/spot-audio-assembly/scripts/sfx_onsets.py audio/sfx/<reel>.wav --cut 0.83:2.10 --out audio/sfx/<name>.wav
```

- **VO never overlaps an sfx**; the audio signature sits on the END CARD and its hit lands on the card's
  burst (`card + 0.02`).
- Music: the **AceDataCloud Suno API is the default route** — async, two takes a call, behind the cost
  line and a GO; the operator's own candidates and a library bed are FALLBACKS that need the operator's
  permission first. A cue is CUT to the marker it ends on, the file name carries the length; a cue
  SHORTER than its span is looped on its own beat grid by whole bars, the take's ending kept
  (`scripts/music_loop.py`), before a longer cue is priced — a cue neither looped nor long enough just ends
  mid-shot and `edl_check` fails it; the earlier accepted cue wins; ducks under speech with ramps;
  Opus-in-`.m4a` is transcoded to 48 kHz WAV before anything reads it. Designed-sound extras (chimes)
  are opt-in, default off.

Sourcing, licences, the signature's geometry, cue and duck rules: [`references/SFX-AND-MUSIC.md`](references/SFX-AND-MUSIC.md).

**Done when:** every sfx and cue is a 48 kHz WAV under `audio/`, its licence or receipt is noted, the
syllable map was confirmed by the operator, and no cue is shorter than its span.

## 4. Clean the native audio — locate, mute, fill

Every take's audio head is scanned before it is mixed; the ear decides which findings are artifacts.

```bash
python3 ~/.claude/skills/spot-audio-assembly/scripts/audio_head_scan.py --root <project> --edl edit/<SPOT>-EDL.json
python3 ~/.claude/skills/spot-audio-assembly/scripts/roomtone_synth.py --ref takes/<sibling>.mp4 --ss 3.2 --t 0.6 --dur 2.5 --out audio/sfx/<SPOT>-roomtone-synth.wav
```

A generated tone or drone (a sustained note under a quiet cut doubles a quiet bed and reads as a musical
artifact), a click where a cut truncates a sound, a line a reused clip must not carry: the
window is muted (`native_audio_from/to` in the EDL) and the hole is filled with **synthesized** room tone
from a clean slice of a sibling take — never a pasted slice (it carries the take's artifacts), never
digital silence (the drop from room fuzz to nothing is audible).

A recorded VOICE — documentary chant, an interview, a testimonial — gets three more rules, each the price of a build the
operator sent back (an auction spot, 2026-09-15):

- **A dip or a mute is placed by the level envelope, never by a transcript gap.** On fast speech the ASR drops words, so a
  "silent" gap in the transcript can be the speaker at full level; a −18 dB dip placed on one cut the caller's own calling.
  Read the RMS envelope of the window first; mute only where it shows no speech.
- **Processed voice enters a build only after the operator's A/B reel.** Light broadband denoise (`afftdn` at about 9 dB)
  is the default and needs no reel; a model separation (DeepFilterNet, demucs) is for a NAMED intrusion — a truck under one
  line — and even then the reel decides: the operator hears "too much messing with his voice" on every model pass a build
  ships unheard. `scripts/voice_ab_reel.py` builds the reel (each treatment once, a gap between, the index printed and
  written beside it); the pick goes into the EDL by ordinal.
- **A tick train under speech is swapped, not patched.** A spectral patch clears one click; eight ticks across half a
  second of speech survive every local repair. Replace the moment — audio AND picture — with another take of the same beat.
- **Speech over a vehicle, a crowd or wind: a speech-enhancement model first, a music/voice separator second, spectral
  denoise third.** Measured on one 5.6 s line as speech-window over pause-window SNR in 300–4 kHz: the original 2.8 dB,
  demucs vocals 5.9, vocals + `afftdn` 8.3, DeepFilterNet3 14.3 with the words intact (whisper p 0.90); ffmpeg
  `dialoguenhance` broke the words. Clean the WHOLE take segment, never the excerpt, so the background does not return at
  the next cut inside it; then re-measure the cleaned excerpt — it reads quieter, and the stem's gain clamp must admit it.
  The order says where to start; the operator's reel still decides what ships.

**Done when:** every flagged window has been listened to, each artifact has a mute + room tone in the
EDL, the scan is clean on the takes the cut uses, and no processed voice is in the stem without a reel pick behind it.

## 5. Build the stem and the captions

```bash
python3 ~/.claude/skills/spot-audio-assembly/scripts/vo_word_times.py --root <project> --edl edit/<SPOT>-EDL.json [--engine local] [--only L2]
python3 ~/.claude/skills/spot-audio-assembly/scripts/build_vo_stem.py --root <project> --edl edit/<SPOT>-EDL.json
python3 ~/.claude/skills/spot-audio-assembly/scripts/build_captions.py --root <project> --edl edit/<SPOT>-EDL.json
```

- The stem places **every** entry with a file (an initial-letter filter once dropped the O.S. line),
  gained to the target with a peak cap; the finisher rebuilds it on every master.
- **Every line's source is read before it is placed**: the builder prints its `source` (a TTS job id, a
  clone id, `extracted_from: <path> @ <s>`) and its gap floor, and refuses a line whose quietest tenth sits
  above `--floor-max` (−40 dBFS by default, measured on one job; a recorded voice or another TTS vendor is measured
  first, and that measurement sets the project's value). A VO excerpted from a finished cut carries that cut's bed
  and scene audio inside every line, and every placement instrument validates it against itself (`floor_ok: "<why>"`
  admits a line that
  legitimately carries sound under it). [`references/VO-PIPELINE.md`](references/VO-PIPELINE.md) § provenance.
- Word times: local whisper is free; Scribe is billed per minute (cost line). **`--only` when one take
  is swapped** — a full re-run clobbers hand patches on the other lines; a swapped file on a lip-synced
  line also re-measures its `at` (`video-edit-edl`). **Transcribe the WHOLE source, then window** — a short
  excerpt transcribed alone lost the agreement between two model sizes that the full source restored.
  **Fast speech — chant, rap, an auctioneer, an overlapping crowd — is beyond ASR onsets**: seven runs (three
  models, three windows, 0.8× speed, a band-limited envelope) put one word's onset anywhere in a 0.74 s
  spread. A cue that must sit ON such a word takes its time from the operator (tapping along, or a tolerance
  stated up front), or from an envelope onset the operator confirmed by ear — never from ASR word times
  alone, and a ±2-frame standard is not offered on it.
- Captions: the narrator's lines only, **never punctuation**, a card ends a frame before the next card's
  lead, nothing over the end card. The display text is the SCRIPT's (`vo[line].text`) and the timing the
  transcript's — the word counts must agree. The style is the campaign's `caption_style`: the house default
  (ALL CAPS, the spoken word in the brand colour) or a variant set by measuring the client's reference
  (`highlight: none`, sentence case, a weight, a line pitch); an unknown key stops the build, and the
  block's placement is checked against our own frames, never over a face.
  Standard: [`references/CAPTIONS.md`](references/CAPTIONS.md).

**Done when:** the stem's printout shows every line at its EDL time under the peak cap with a clean floor
and a recorded source, and the caption manifest's cards match the EDL's word ranges and the script's words
with no punctuation.

## 6. Read the delivered file

After `video-finish-qc` renders, the audio is verified on the DELIVERED file, per metric with its label —
never a total score:

```bash
python3 ~/.claude/skills/spot-audio-assembly/scripts/qc_vo_placement.py --root <project> --edl edit/<SPOT>-EDL.json --deliv deliver/<file>.mp4
ffmpeg -v info -i deliver/<file>.mp4 -af ebur128=peak=true -f null - 2>&1 | tail -12      # I · LRA · true peak
```

- Placement by **envelope** correlation (10 ms RMS, 300–4 kHz): a waveform correlation reads 0.08 under
  loudnorm + AAC where the envelope reads 0.65–0.91; every line < 15 ms or it is OFF. The instrument
  self-tests first or prints nothing.
- Loudness and true peak against the EDL's targets: the delivered file is judged against the PLATFORM ceiling
  (`loudnorm.TP_ceiling`, −1 dBTP unless the platform says otherwise), never against the master's TP plus a
  fixed allowance — the AAC overshoot on limited peaks measured 0.4–1.0 dB and grows with how hard the limiter
  works (a quieter premix that pushed it harder overshot by 0.5 dB where earlier versions overshot 0.2), so the
  master's TP sits under the ceiling by at least that and the gap is re-measured after any premix change;
  measured per channel, never on a mono sum.
- **No operator in the loop (a headless run) makes every pick and listen PROVISIONAL**: the alternatives stay on
  disk by ordinal and path, the instruments stand in for the ear, and a listen queue — timecodes and what to
  listen for — rides the handoff; nothing is final until the operator has done the queue.
- The VO ↔ sfx clash and the head tones are listened for at every hit, the signature and every cut.

The mix graph the finisher implements (a static sum — no loudness processing in the master), the loudness
pass (one measured static gain, a limiter at 192 kHz before the resample and again after it, the limiting the
target costs printed; the client reference's own level as the target when there is one), and the two reads for
"the music changes around the voice": [`references/MIX-AND-QC.md`](references/MIX-AND-QC.md).

**Done when:** `PLACEMENT OK`, I within 0.5 LU and TP under the EDL's ceiling on the delivered file,
and the listen at each hit found no clash.

## Failure behavior

- No `ELEVENLABS_API_KEY` in the environment → the script stops and says to load the env file that holds it; the key is
  never pasted on a command line or into a job file.
- An HTTP 401 → probe the key with `curl …/v1/user` before anything else; the CLI's own 401 text always
  blames the keyring.
- A 429/5xx → three retries with backoff, then stop; a refused line (moderation) → the dub procedure,
  never a rephrase the client did not write.
- A self-test failure in any instrument → the instrument is fixed before any number from it is used.
- `VO SOURCE HYGIENE FAIL` → find the line's clean source (the TTS job, the clone render, the client's dry file); never
  raise `--floor-max` to pass an excerpt of a finished mix. `floor_ok: "<why>"` admits only a line that legitimately
  carries sound under it, and its why is the record.
- A caption build that stops — an unknown `caption_style` key, a script line whose word count differs from its word
  times — is fixed in the style block or by `vo_word_times.py --only <line>`, never by hand-editing the cards.
- A placement OFF, a TP over the ceiling, a clash under the signature → back to the EDL, re-render; never a
  gain nudge on the delivered file.

## Scripts

| script | does |
|---|---|
| `gen_vo.py --root --jobs [--voice] [--dry-run] [--preview] [--door]` | TTS takes by REST; characters before/after; skip-by-name; reel + index; door treatment |
| `vo_word_times.py --root --edl [--engine local\|scribe] [--only]` | word times per line; `--only` preserves hand patches |
| `build_vo_stem.py --root --edl [--out] [--floor-max]` | every placed line on the stem at target LUFS with a peak cap; each line's gap floor (refuses a line carrying a mix) and its source |
| `build_captions.py --root --edl [--out-dir] [--fonts-dir]` | cards at delivery resolution + manifest; the script's words on the transcript's times; every style key checked; `highlight: none` = one plain layer per card |
| `sfx_onsets.py <file> [--cut a:b --out]` | the syllable map by ordinal; sample-exact cuts with fades |
| `audio_head_scan.py --root (--edl \| --files)` | stable tones/drones in each take's head, labelled |
| `roomtone_synth.py --ref --ss --t --dur --out` | stationary room tone coloured by a clean slice |
| `voice_ab_reel.py --out <reel.wav> name=<file> … [--gap 0.3]` | the A/B reel of voice treatments — each variant once, a gap between, the index printed and written beside the reel; `--selftest` |
| `music_loop.py --take --bars --a-target [--bpm --first-beat] [--xfade] [--out]` | a short bed looped on its own beat grid by whole bars, the take's ending kept and its hit moved by exactly the jump; a new file, never an overwrite; `--selftest` |
| `qc_vo_placement.py --root --edl --deliv` | envelope-NCC placement of every line on the delivered file |

## Cross-references

- [`references/VO-PIPELINE.md`](references/VO-PIPELINE.md) · [`references/SFX-AND-MUSIC.md`](references/SFX-AND-MUSIC.md) ·
  [`references/CAPTIONS.md`](references/CAPTIONS.md) · [`references/MIX-AND-QC.md`](references/MIX-AND-QC.md).
- `video-gen-cost-gate` — the cost-line and GO form every billed call uses.
- `video-edit-edl` — where each sound sits and the gate that holds it; `video-finish-qc` — the render that
  mixes it; `mastering-audio` — a standalone loudness master.
- Handoffs follow `~/.claude/skills/video-production/references/CHAIN.md`.
- Deeper context — the upstream projects behind the rules here: `video-production/references/CONTEXT-MAP.md` § Where the deeper context lives.
