submit-workflow · git:20260830.7e83781 · 2026-08-30 · sha256 8e0d13d8b5ab920a

submit-workflow git:20260830.7e83781A

Immutable. This exact content is served forever at /api/v1/blob/8e0d13d8b5ab920a.

---
name: submit-workflow
description: Use when submitting, validating, or debugging NPA SkyPilot workflow YAMLs and workflow runner paths.
---

# Submit Workflow

## When To Use

Use this skill for workflow launch, YAML validation, runner scripts, and
SkyPilot submission behavior.

## Procedure

1. Read `skills/tools/skypilot-workflows/SKILL.md` for SkyPilot version and
   cleanup constraints.
2. Prefer `npa.workflow/v0.0.1` specs under
   `npa/workflows/workbench/npa-workflows/`. Parse / `validate-spec` locally
   before launch.
3. Use `NPA_SKYPILOT_BIN` or `npa skypilot status --bin-path`; do not assume
   `sky` from `PATH`.
4. Submit through `npa workbench workflow submit` (accepts npa.workflow specs
   and legacy SkyPilot YAML) or the shared workflow submission helper.
5. Keep cleanup best-effort and avoid tearing down a shared controller unless
   the operator explicitly requests it.

## Three-Tier Contract

- CLI: `npa workbench workflow --help` and tool-specific `workflow` commands.
- SDK: use shared workflow submission helpers rather than shelling out from
  application logic.
- YAML: author shipped workflows as `npa.workflow/v0.0.1` specs under
  `npa/workflows/workbench/npa-workflows/`. `npa workbench workflow submit`
  accepts those specs (plans, renders, then launches SkyPilot) and still accepts
  raw SkyPilot YAML supplied by an operator or by guarded single-task example
  directories.

## Live submit prerequisites (real cluster)

A real `npa workbench workflow submit` (not `--plan-only`) needs, on top of a
successful `npa skypilot verify --cluster <exact-context>`:

- **Secrets via `--secret-env`** (never in the YAML): `NEBIUS_TOKEN_FACTORY_KEY`
  for Token Factory stages, `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` for S3,
  `HF_TOKEN` / `NGC_API_KEY` for gated model pulls. Submit resolves each requested
  name from the explicit process environment first, then the selected project's
  configured NPA credentials; it fails locally if the value is unavailable.
- **`NPA_SRC_S3_URI` (or `--image`)** for CPU tool steps and `run.shell` states —
  they have no heavy workbench image and install npa from that source tarball,
  else render fails with "planned step has no workbench image and NPA_SRC_S3_URI
  is unset". Persist it once with `npa configure --src-s3-uri s3://bucket/prefix/npa`
  so a new shell resolves it from `~/.npa/config.yaml` instead of failing preflight
  on already-staged objects (`scripts/stage-npa-src.sh` does this for you).
- **`--assume-decision promote_checkpoint`** for specs with a dynamic gate/loop.
- **`--var key=value`** to override `config` (e.g. `--var bucket=<real-bucket>`;
  the reference specs default to `bucket: example-bucket`).

## Gotchas

- **Kubernetes controller launch is transactional.** Immediately before every
  Kubernetes managed-job launch, NPA probes `/readyz` on the exact selected
  context with SkyPilot's `KUBECONFIG` environment. Readiness requires three
  consecutive successes spanning 10 seconds. After any failed/uncertain launch,
  NPA reconciles exact `sky jobs queue --all --output json` evidence: adopt one
  immutable ID, retry only after authoritative absence plus a classified
  transport/API warm-up failure, or fail closed as indeterminate. Never bypass
  this with raw `sky jobs launch`, retry by name, or cancel by name.
- **An isolated SkyPilot state directory has its own stable controller user
  identity.** Reuse the same directory when resuming a run; a different isolated
  directory intentionally selects a different controller namespace. An explicit
  `SKYPILOT_USER_ID` still takes precedence.
- **A baked workflow image validates the module it actually executes.** Set
  `config.baked_npa_import` to that dotted module when `require_baked_npa` is
  enabled; otherwise the backward-compatible probe is `npa.cli.main`. This keeps
  source attestation strict without requiring narrow stage images to install the
  unrelated full CLI dependency closure.
- **Baked Kubernetes tasks get writable bootstrap caches.** NPA supplies
  pod-local `XDG_CACHE_HOME` and `UV_CACHE_DIR` defaults under `/tmp` so a
  read-only image-owned model cache cannot break SkyPilot's setup probe. Explicit
  workflow environment values still take precedence; model/checkpoint caches and
  mounted durable volumes are not redirected.
- **Explicit workload retries apply after exact resume reconciliation too.** If an
  adopted in-flight job is proven terminal, `--retries` advances through the same
  durable terminal-retry path and assigns a new attempt identity. With no explicit
  retries, the terminal outcome remains preserved and no duplicate is launched.
- **Infrastructure recovery has its own finite policy.** `--retries` remains the
  payload/terminal-wave retry count. `--max-infrastructure-recoveries` bounds
  typed capacity, quota, node-not-ready, and provider recovery per wave (default
  1; 0 disables automatic relaunch). Exhaustion is persisted and terminal; the
  two policies never silently borrow from each other.
- **Runtime supervision is durable and fail closed.** Pending pods are inspected
  by exact managed-job ID. Image/auth/reference, missing Secret/ConfigMap,
  malformed pod config, and impossible GPU shape failures stop immediately and
  cancel only that ID. Proven transient infrastructure failures may create a new
  immutable attempt under the same run ID only after immutable workflow/source/
  image identity, declared S3 output absence, preflight readiness, and exact
  cancellation are verified. Expected identities are independently recomputed
  from the current spec, source selection, and digest pins rather than copied
  from the attempt being checked. Unknown evidence blocks relaunch.
- **Async acceptance is not workload observability.** SkyPilot launch uses its
  asynchronous API mode, then the existing launch transaction reconciles the
  exact logical name to a provider job ID before runtime polling begins. Exact
  cancellation is polled to terminal; a request acknowledgement alone never
  permits relaunch.
- **Checkpoint recovery is capability-based.** Completed waves require validated
  declared outputs. Mid-stage resume requires an explicit compatible loader and
  validated application checkpoint; otherwise recovery restarts the incomplete
  wave and must not claim checkpoint resume. The same adapter contract is active
  in `npa workbench genesis train-teacher --runtime serverless` for deterministic
  Nebius Serverless Job re-attempts without a GPU supervisor VM. It does not
  enable mixed per-stage Serverless routing in `npa.workflow/v0.0.1`.
- Transaction recovery uses capped exponential jitter and a 180-second recovery
  deadline. This is product behavior, not an operator job/time budget. A
  recovered launch proceeds in the same command; use `--resume-run <same-id>`
  only for crash/restart or a printed indeterminate/deadline recovery action.
- A failed reconciliation with launch sequence zero created no SkyPilot job.
  NPA records a completed no-op rollback instead of leaving a
  `recovery-required` journal that would block unrelated project operations.
  Any failure after a launch may have been issued remains recovery-required.

- SkyPilot `envs` does not support self-referencing interpolation. The
  npa.workflow renderer resolves images and config before submit so rendered
  YAML has no `${VAR}` placeholders.
- `sky jobs launch` does not provide a reliable dry-run path in the pinned
  version; use `npa workbench workflow submit --plan-only` for npa.workflow
  specs, or mock submission before live launch.
- Mixed serial and parallel task groups can be fragile; serialize when behavior
  must be deterministic. Parallel sweeps stay SkyPilot-only in v0.0.1.
- **GPU accelerator name is cluster-specific.** Specs use canonical
  `RTXPRO6000:1`, but a cluster may only advertise the raw label (e.g.
  `RTXPRO-6000-BLACKWELL-SERVER-EDITION`), and the name changes while the NVIDIA
  GPU operator is still labelling nodes (`nebius.com/gpu-name: RTX6000` first,
  `nvidia.com/gpu.product` after). A mismatch fails with `FAILED_PRECHECKS` /
  "cluster does not contain any instances satisfying the request" — not a capacity
  problem. Submit now remaps this automatically; use
  `npa workbench workflow gpus --cluster <name>` to see the names yourself, or
  `--no-resolve-accelerators` to submit the spec's values verbatim.
- **`NAME:N` needs N GPUs on one node.** SkyPilot places all GPUs of a task on a
  single node, so `NAME:2` can never schedule on 2 nodes × 1 GPU no matter how many
  nodes exist. `workflow gpus` prints the requestable quantity per node; submit
  rejects anything above it. Multi-GPU fan-out docs assume N GPUs per pod, which is
  a different cluster shape from "N single-GPU node presets".
- **A workflow's images resolve from GHCR releases unless explicitly overridden.**
  `npa configure` records the public release namespace by default. Run
  `npa workbench workflow preflight-images <spec.yaml>` — it reports each image as
  `ok`/`not_found`/`forbidden` and prints the build command for the tag
  `npa/src/npa/deploy/images.py` pins (the guide's tags are pinned to those by
  `tests/guardrails/test_paidf_image_tags_match_code.py`). `submit` runs the same
  check **before `deployIfAbsent`**, so a missing release costs no
  cluster time.
- **Multi-tool validation images stay distinct.** Repeat
  `--image-override TOOL_REF=IMAGE` on preflight and submit. Exact tool refs take
  precedence over the optional global `--image`; preflight resolves each selected
  artifact to the digest the renderer uses.
- **A registry `403` stalls rather than fails.** Kubernetes retries image pulls
  forever, so an unpullable image leaves the job in `PENDING`/`ImagePullBackOff`.
  Listing a repository's tags is a *different permission* from pulling it, so a
  `200` on `/v2/<repo>/tags/list` proves nothing. Submit reproduces each planned
  pull with the credentials it injects and refuses to launch on a `403`; run it
  standalone with `npa workbench workflow preflight-images <spec.yaml>`, or skip
  with `--no-preflight-images`.
- **A large authenticated cold pull is not an access failure.** Bootstrap probes
  default to a 30-minute observation window. Use
  `--image-bootstrap-timeout-seconds 0` for no deadline while warming large
  images; digest, authentication, attestation, capability, exact ownership, and
  verified cleanup gates remain mandatory.
- **A silent 15-minute submit is usually the kubernetes client.** SkyPilot 0.12.2
  does not cap the client version, and client 36+ makes every `pod_config` fail
  validation, so the managed-jobs controller retries forever. `npa skypilot
  bootstrap` pins a working client and repairs an existing venv; `npa skypilot
  status` reports the installed version. Submit streams SkyPilot output live and
  names this failure when it appears.
- **Stale `NEBIUS_IAM_TOKEN` breaks sky/terraform.** The Nebius provider prefers
  an ambient (often expired) `NEBIUS_IAM_TOKEN` over the fresh CLI token, giving
  `PermissionDenied` / `Unauthenticated` even though the `nebius` CLI works.
  `unset NEBIUS_IAM_TOKEN NPA_NEBIUS_IAM_TOKEN` before submitting/deploying.

## Teardown

- **Cancel then wait, then tear down.** Use `npa workbench workflow cancel
  <run-id> --project <alias> --json`; a planned/staged run that never launched is
  a successful repeat-safe no-op, while a launched run uses NPA's pinned
  SkyPilot runtime and waits for the exact manifest-proven job. Only after all
  workflows are terminal, remove the shared controller with `npa skypilot
  cleanup-controller --yes`. The underlying helpers retry the specific
  in-progress-jobs refusal after the queue drains.
- **A PENDING job may be dead, not slow.** A pod stuck in `ImagePullBackOff` or
  `Unschedulable` is retried by Kubernetes forever, so the job never becomes FAILED.
  `npa workbench workflow status` reports the pod-level reason for a PENDING job.
- **`npa cleanup`** reports what a teardown left behind (local caches, project
  entries, non-terminal managed jobs, the service accounts `npa configure` creates)
  and prints the ordered runbook. `--yes` removes the local caches only; it never
  deletes cloud resources or service accounts.
- **`npa cluster down`** previews the PodDisruptionBudgets that will hold up the
  node drain, so a multi-minute silence is expected rather than alarming.

## Verify

```bash
npa/.venv/bin/python -m pytest npa/tests/guardrails/test_skills_index.py -q
```

The smoke test invokes workflow help and parses representative workflow YAML.