---
name: groot
description: Use when working on NVIDIA GR00T deployment, model download, finetuning, evaluation, serving, inference, conversion, status checks, validation, routing, or CUDA alignment.
---

# GR00T

## When To Use

Use this skill for NVIDIA GR00T robot foundation model workbench changes,
especially NGC/Hugging Face model handling, embodiment tags, checkpoint
conversion, serving, inference, and validation.

## Procedure

1. Start with the current command surface:

   ```bash
   npa workbench groot --help
   ```

2. Use `download` to stage model artifacts and credentials, `finetune` for
   training, `eval` for offline scoring, `serve` and `infer` for runtime calls,
   and `convert` when transforming checkpoints for downstream use.
   For declarative N1.7 training, use
   `workflows/testing/groot-1-7-finetune.yaml`; its
   `gpu_count` is a positive single-node world size that controls both the
   scheduler allocation and the real trainer.
3. Use `status`, `system-info`, and `list` for operational checks. Keep
   `ensure-ingress`, `register-byovm`, `reload-env`, and `cleanup-partial`
   scoped to setup and recovery flows.
4. Preserve credential redaction for NGC, Hugging Face, S3, and SSH values.
5. Before deploy provisions or updates anything, validate actual Hugging Face
   access to both the selected GR00T checkpoint and its runtime-fetched
   `nvidia/Cosmos-Reason2-2B` dependency. The operator's HF token and upstream
   permissions are the only local gate for gated weights; do not add a manual
   acceptance flag or a model-check bypass.

## Three-Tier Contract

- CLI: `list`, `deploy`, `download`, `finetune`, `eval`, `serve`, `infer`,
  `convert`, `status`, and `system-info` are the main user commands.
- SDK/API: keep model, checkpoint, and storage path normalization in shared
  helpers so service and CLI routes do not diverge.
- YAML: `groot-1-7-finetune.yaml` calls `workbench.groot.finetune` in the
  stage's own image. Keep model/source pins, checkpoint S3 URIs, run ID, and
  GPU count explicit; counts above one must use the upstream `torchrun` path.

## Routing And Validation

- GR00T does not require RT cores for the standard model paths.
- Route throughput-heavy training/eval to H100/H200 unless a command or image
  specifically requires another target.
- CUDA 13 alignment is vendor-paced on NVIDIA x86_64 CUDA 13 and is not a
  Nebius infrastructure blocker.

## Runtime Isaac bootstrap (the container ships no Isaac Sim)

Before a GR00T Isaac simulation run, load
`skills/atomic/third-party-eula-preflight/SKILL.md`. Standard inference and
fine-tuning do not trigger this preflight.

The `npa-groot` image contains **no NVIDIA Isaac Sim or Isaac Lab code**. It used to bake
Omniverse Kit, which made it non-redistributable; Isaac is now downloaded on first use of
`/isaac-sim/python.sh` from `https://pypi.nvidia.com`, into a cache volume, under the
**operator's own EULA acceptance**. Full rationale:
`docs/workbench/container-packaging.md` and `skills/atomic/solution-licensing/SKILL.md`.

What this changes in practice:

- **Only Isaac simulation defaults acceptance.** NPA defaults
  `ACCEPT_EULA=Y` on that path. Empty, `N`, `NO`, `0`, `FALSE`, or
  `--no-accept-eula` exits **78** before download. `Y`, `YES`, `1`, and `TRUE`
  are accepted case-insensitively; other values are invalid. The launcher derives
  `OMNI_KIT_ACCEPT_EULA=YES` internally; do not expose duplicate user plumbing.
  Keep `PRIVACY_CONSENT` and telemetry off. Standard GR00T inference and
  fine-tuning do not require Isaac acceptance.
- **Reach Isaac through `/isaac-sim/python.sh`** (the value of `ISAAC_LAB_PYTHON`). That is
  the bootstrap shim, and it is what every SkyPilot template, the sim2real engine and the
  workbench CLI already use. A bare `python3` is the *system* interpreter and will not
  find Isaac.
- **Never invoke the shim from a Dockerfile `RUN`.** It would download and bake ~4.5 GB of
  Isaac into a layer. Build-time work uses the image's own venv python.
- **Budget the first start.** Measured on RTX PRO 6000: 111 s cold, 32 ms warm, 10.04 GiB
  of cache. Pre-warm a shared volume once per node/PVC with
  `npa/docker/workbench/common/warm-isaac-cache.yaml`, then run workload pods with
  `NPA_ISAAC_CACHE_READONLY=1`. Otherwise every pod pays it, and 8 GPU pods on a node
  download ~36 GB.
- `isaac-bootstrap status` reports what is cached without needing acceptance or network;
  `isaac-bootstrap verify` additionally launches Isaac Sim headless (needs a GPU).
- No NGC credentials are needed to build or run this image.

GR00T inference and fine-tuning run in `GROOT_VENV`
(`python -m npa.smoke.test_groot_functional` covers inference) and need no Isaac
or EULA acceptance, so they pay no first-run download. Only the Isaac Lab
simulation paths do.

## Gotchas

- NGC credentials are required for gated NGC model refs. Do not print token
  values in diagnostics.
- Managed VM `deploy` defaults to in-place updates for existing aliases.
  Terraform plans that would destroy or replace critical infrastructure are
  blocked unless the operator passes `--replace` and confirms with `--yes`.
- BYOVM deploys record `endpoint_strategy: public` or
  `endpoint_strategy: ssh_fallback` in `~/.npa/config.yaml`; live commands honor
  that strategy and can self-heal blocked public endpoints through a transient
  SSH-local route.
- Known issue: output truncation at high step counts must be validated with
  artifacts, not subjective evaluation.
- Multi-GPU fine-tuning defaults to NCCL's native transport selection. When a
  clean two-rank collective proves that both P2P and SHM are unsafe on a
  single-node host, use the workflow's `nccl_transport=socket` compatibility
  fallback. It disables both transports, so do not use it speculatively on
  healthy high-bandwidth hosts.

## Verify

```bash
npa/.venv/bin/python -m pytest npa/tests/guardrails/test_skills_index.py -q
npa/.venv/bin/python -m pytest npa/tests/orchestration/npa_workflow/test_groot_finetune_workflow.py -q
```

The skill smoke invokes current GR00T training help and parses the reference
N1.7 workflow in addition to the download/convert command checks.
