---
name: shot-production
description: Produce image, video, or audio media for an approved historical-drama Shot. Use when combining Shot and Scene context, selected stable Assets, and prior Media into start frames, end frames, video, dialogue, ambience, or other physical media outputs.
---

# Shot Production

Use `shot.get_shot` and `scene.get_scene` when their stable IDs are known and the approved production context was not already supplied. For a continuous Scene, establish one shared Sequence Context before producing its Shots. Load [references/production-rules.md](references/production-rules.md) to plan stable facts, references, Shot deltas, review, and revision. When a Shot has `spokenContentBindings`, also load the [Dialogue Layer content convention](../../docs/dialogue-layer-content-convention.md), resolve every binding against the Scene's canonical `spokenContent`, and require numeric `plannedDurationMs` plus a passed `DURATION_FEASIBILITY` check before physical media production. Keep these decisions in Agent Run Context rather than creating new domain records.

Discover reference candidates from named visible characters, the Scene, focal props, costume variants, and other relevant stable visual entities. Use `asset.get_asset`, `asset.list_assets`, or `asset.search_assets` without inventing references. Inspect stable reference Media with `media.get_media`; use `media.list_media` only for a clear structural media scope, not broad discovery. Select no more than three stable Asset-plus-Media references. Record selected and omitted candidates with rationale. Return `MISSING_STABLE_REFERENCE` for a key visible character whose stable identity is absent; do not silently omit it, substitute an informal image, or generate an unreviewed master inside Shot production. This blocks only media that must show that character; it does not invalidate the Work-scoped `speakerKey`, Scene dialogue authoring, narration, or non-visual speaker identity.

Do not require a visual provider for context reads, non-visual planning, Shot design, or other non-visual work. For image or video planning and execution, load [references/visual-provider.md](references/visual-provider.md) so the plan includes the complete reference-to-provider-to-Media handoff. Before actual execution, preflight only the Drama and visual capabilities required by that request. Return `DRAMA_PROVIDER_UNAVAILABLE`, `VISUAL_PROVIDER_UNAVAILABLE`, or `VISUAL_PROVIDER_CAPABILITY_MISSING` when the corresponding capability is unavailable; stop rather than installing, configuring, or simulating a provider.

Before preparing new video inputs, consume a frozen decision from [video model selection](../video-model-selection/SKILL.md). Select only after Shot narrative, performance and sound requirements are fixed. The selected mode determines which existing inputs to reuse and which are genuinely missing. For video and its necessary input images, use one stage budget in [first-pass production](references/first-pass-production.md); submit only the reserved decision request and requalify after any material change.

Before visual generation, compile each approved Shot into stable identity and environment facts plus current action, composition, representative keyframe intent, required visual evidence, forbidden visual outcomes, and continuity constraints. For new image production and image revisions, load [first-pass production](references/first-pass-production.md) and use its executable request preflight and persistent reservation/review gate before calling the Host provider. Qualify up to three representative planned Shots before expansion, preserve reference versions and actor/action ownership, and pause on repeated major failures. Use resolved dialogue only for the visible delivery, reaction, off-screen, or voice-over intent relevant to the binding; never copy its text into the Shot or provider-owned state. Resolve each selected input through `media.resolve_media`, execute the available visual capability, and apply the per-Shot review defined in the production rules. On failure, allow at most one targeted revision while preserving Stable Facts and the Reference Plan unless the review proves that plan incorrect. Do not revise a Shot that already passes.

When an authoritative `DPDSnapshot` is supplied for a visible speaker, first project it into a typed, provider-neutral `VisualPerformanceBrief`. The brief owns only current visible behavior such as body activity, head behavior, gaze, facial tension, gesture policy, interaction orientation, pre-speech behavior, and visible control. Stable Character/Scene appearance remains in Asset/Media; framing, lens, angle, composition, and camera movement remain in Shot design. The projection must not reinterpret objective, tactic, relationship, subtext, or historical context. Materialize the brief together with the separate Shot camera/design facts, and record unsupported or approximate Provider controls without pretending they are exact.

After technical/source checks, preserve a worthwhile paid output as CANDIDATE even while content review is pending. Use the production gate `persist` operation for an existing visual attempt: call `media.import_media` or reuse by stable source, get/resolve/download, verify full bytes and ownership, and write/read back its formal business binding. Keep content review and user adoption independent. Any required identity annotation is a separate derived version; it is not a Provider quality criterion and must not replace already adopted bytes. Keep the verified `mediaId` and role in both the formal business record and Agent Run Context. For a DPD-directed performance video, inspect playback or a controlled set of representative frames and record an accepted `RealizedPerformanceSnapshot` from actual visible facts, including only the minimum useful timing windows. Describe the video even when it deviates from DPD; a conformance diagnostic may inform regeneration but must not block Media, Snapshot, or downstream work and must never rewrite the observation. Unknown or obscured facts remain `UNKNOWN`. Observation never infers objective, tactic, relationship, subtext, or internal activation. Compare only Review-PASS Shots during cross-Shot continuity review. If two or more Shots require continuity regeneration, stop with `SEQUENCE_CONTINUITY_REQUIRES_REPLAN`. When audio production is explicitly requested, `production.generate_role_dubbing` may consume the resolved Scene text, Work-scoped speaker identity, performance intent, provenance, estimated duration, and Shot coverage intent; it must not rewrite, delete, split, merge, or replace canonical `spokenContent`. There is no mandatory image-to-video-to-audio sequence and no separate Dialogue record or Audio timeline in this Skill. An explicitly adopted existing AV may go directly to Host assembly preparation with its original sound; source dialogue bindings are not a mandate to add external speech to an already adopted performance. Record a material narrative gap separately without silently rewriting the screenplay or automatically scheduling dubbing.

Use `context.build_context` only when required Shot context was not supplied, and use `context.refresh_context` only after stable generated Media makes the current context stale. Never expose temporary URLs, storage locations, filenames, provider task IDs, workflow documents, node IDs, or provider-specific responses as Drama domain facts. For a DPD-directed performance video, stop when both the stable Media role and accepted Realized Performance Snapshot are clear; otherwise stop when the requested media exists in stable Drama memory and its role is clear. Do not redesign the Shot or continue to another Skill automatically.

## Dialogue-coupled execution

When complete actual Audio exists, derive this Video's execution timing from its measured durations, the immutable DialogueTimingPlan's protected reactions/holds, and the production target. Never let original estimates exclusively determine speech phases. Consume the derived material in VisualPerformanceBrief and the real request; emit active speaker, listener, visible action, reaction, transition purpose and performance boundaries rather than only hashing them.

Keep monolithic production for the current single-Shot baseline. New Video requires new shot-level RP and speaker-specific snapshots with observedSpeakerKey; null denotes aggregate only. Observe visible participation and handoff, not guessed psychological states or mouth-as-speech-onset. A physical fit does not prove visual fit. Respect latest user rejection and the current task's explicit intermediate-review authorization. Record one shared corrective visual rebuild budget; if that fails, do not enter an unbounded Audio/Video loop. Mouth derivatives require fresh face/non-speaker/identity/continuity observation, preserving original source Media and timing authority.


## Durable completion and recovery

For retained image, video, audio or derived AV, formal completion requires full
readback SHA-256/size, decodable media, correct business ownership and a queryable
stable business binding. A local path, upload name, returned ID or temporary URL
alone is PERSISTENCE_PENDING. Content PASS, user adoption and persistence are
independent; persistence must never rewrite an existing adoption decision.

The shared completion entry `complete_retained_media` implements import/reuse,
readback and binding. Visual attempts call it through `visual_preflight.py persist`;
role speech calls it before returning a production result; an explicit mux calls
it on the new output before formal delivery. Adopted native AV consumption uses
`prepare_bound_media(adopted=True)` to recover the selected original from its formal
business reference. It performs no dubbing, separation, mixing or remuxing.

After a successful generation, resume at import/readback/binding with the same
job, output hash and sourceRef. After an uncertain import, query the exact stable
source first; ambiguity, permission denial, hash conflict or unknown ownership
blocks completion. Retry transient reads with finite backoff and fresh resolve.
Never generate again to repair missing cache or persistence. Do not import
unneeded DEBUG/REJECTED files or invent an output that was never generated.
