vss-benchmark-vlm-qa · v3.3.0 · 2026-09-19 · sha256 d9049af6d1e7a072

vss-benchmark-vlm-qa v3.3.0B

Immutable. This exact content is served forever at /api/v1/blob/d9049af6d1e7a072.

---
name: vss-benchmark-vlm-qa
description: Benchmark video Q&A accuracy and latency of a deployed RT-VLM (Cosmos Reason 3) via vss vlm run, using questions and videos from the DSS vss-devx-base dataset. Replaces the deprecated nat eval / vss-agent QA path. Not for tool-calling or trajectory evaluation, and not for LVS summarization throughput.
license: Apache-2.0
metadata:
  version: "3.3.0"
  author: "NVIDIA Video Search and Summarization Team"
  github-url: "https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization"
  tags: "nvidia blueprint performance benchmarking vlm qa"
---

# Benchmark video Q&A via `vss vlm`

Measure **accuracy** (LLM-as-judge vs ground truth) and **latency** of end-to-end
video question answering by calling **`vss vlm run`** against a deployed Cosmos
Reason 3 RT-VLM. Questions and clips come from DSS dataset **`vss-devx-base`**
(`nvdataset`).

This replaces `docker exec vss-agent nat eval` for the QA slice. It does **not**
score tool-calling or trajectories.

## When to use

- The user asks to benchmark / evaluate VLM video Q&A after vss-agent / NAT eval
  was removed.
- The user wants latency and answer accuracy on `vss-devx-base`.

## When not to use

- Tool-calling or trajectory evaluation — out of scope.
- LVS summarization throughput — use `vss-benchmark-video-summarization`.
- Ad-hoc single questions — use `/vss-ask-video`.

## Prerequisites

- A VSS stack with RT-VLM serving Cosmos Reason 3, and `vss configure` already run
  so `vss configure check` lists `rt_vlm` as `ok` and `vst` as `ok`.

  **Configure with a routable address, not `localhost`.** Clips are addressed as VIOS
  sensors so RT-VLM fetches them by URL; the URL VIOS mints is built from the
  configured origin. A loopback origin mints a loopback URL, which means nothing
  inside the RT-VLM container, so the CLI falls back to inlining the clip as base64
  and the VLM rejects anything large with `HTTP 422 ... content ... valid string`.
  `vss configure --base-url http://<host-ip>:7777` avoids that — `--base-url` is a
  `vss configure` flag, not a benchmark one. `--inline-media` *is* a benchmark flag; it
  forces the old inline behaviour and is only safe for clips under ~10 MB.
- `uv` and this checkout (CLI via `uv run --project libs/vss vss`).
- The `nvdataset` CLI. It is **not** on PyPI, and the index used by the old
  deep-search eval (`urm.nvidia.com/.../sw-ngc-data-platform-pypi`) returns 403.
  Install from the documented read-only index instead — no credentials needed:

  ```bash
  uv tool install --index https://artifactory.pdx.nvidia.com/artifactory/api/pypi/sw-ngc-data-platform-pypi-local/simple nvdataset
  ```

- DSS access, one of:
  - `NVDATASET_API_KEY` — a **Personal Key** from
    [org.ngc.nvidia.com/setup/personal-keys](https://org.ngc.nvidia.com/setup/personal-keys)
    scoped to the service `NVIDIA Dataset Service`, with the NGC org switched to the
    one owning the dataset. This is *not* the global NGC key used by the NGC CLI; a
    global key returns 403. `NVDATASET_NGC_API_KEY` and `NGC_API_KEY` are also read,
    in that order, for backward compatibility only — the run prints the variable it
    picked as `dss credential: <name>`, so check that line if a 403 surprises you.
  - `nvdataset auth login` (Starfleet SSO), which needs no key. Add `--flow device`
    on a remote box with no browser. Group access requires membership in
    `ngc-datasetservice-viewer-<tenant>-<group>` (reader) or `...-user-...` (writer).

  Plus tenancy, which SSO does **not** supply — after `auth login`, `nvdataset auth
  status` still reports `"tenant_id": null` and every call fails with `Did not find
  tenant_id`. The script names no tenant, so set one yourself: export
  `NVDATASET_TENANTID` and `NVDATASET_GROUPID`, or save them once with `nvdataset auth
  context add`. Ask the dataset's owning team for its coordinates. Another dataset
  needs no change to the script.
- An OpenAI-compatible judge LLM: `EVAL_LLM_JUDGE_BASE_URL` and `EVAL_LLM_JUDGE_NAME`,
  authenticated with `EVAL_LLM_JUDGE_API_KEY`. `NGC_API_KEY` is deliberately **not**
  sent to non-NVIDIA judge hosts — it is set for the dataset download and must not
  reach a third party. Any chat-completions endpoint will do; the judge moves absolute
  scores on its own, so hold it fixed across runs you mean to compare, and read
  `judge_model` in `summary.json` before comparing two numbers. `--skip-judge` gives
  latency only.

Bootstrap is in the repo-root [AGENTS.md](../../../AGENTS.md). Do not construct
RT-VLM URLs; `vss vlm run` reads the recorded config.

## Run

```bash
export NVDATASET_API_KEY=<personal-key>            # or: nvdataset auth login [--flow device]
export NVDATASET_TENANTID=<tenant>                 # SSO does not set this; see Prerequisites
export NVDATASET_GROUPID=<group>
export EVAL_LLM_JUDGE_BASE_URL="${LLM_BASE_URL}"   # OpenAI-compat origin, e.g. http://127.0.0.1:8000
export EVAL_LLM_JUDGE_NAME="${LLM_NAME}"

# Optional: already-extracted dataset
# export VSS_EVAL_DATASET=/path/to/vss-devx-base

<repo>/skills/benchmarking/vss-benchmark-vlm-qa/scripts/run_vlm_qa_benchmark.sh \
  --dataset-name vss-devx-base \
  --dataset-file dataset_single_turn.json
```

Both dataset flags are **required** — the script carries no default dataset, so it
never assumes one team's DSS coordinates.

Useful flags (forwarded to `benchmark_vlm_qa.py`):

| Flag | Purpose |
|---|---|
| `--dry-run` | Resolve QA items and video files; no VLM calls |
| `--limit N` | First N QA items (smoke) |
| `--skip-judge` | Latency only |
| `--skip-download` | Use an already-downloaded `vss-devx-base` |
| `--timeout SEC` | Passed through as `vss vlm run --timeout` (default 300) |
| `--num-frames N` | Frame budget (default 20, matching the old RT-VLM agent config) |
| `--model ID` | Override the RT-VLM model `vss configure` recorded |

Outputs under `<dataset>/../../results/vlm_qa/` (or `--output-dir`):

- `summary.json` — mean accuracy, latency mean / p50 / p90 / p95 / p99, and the
  model the deployment reported serving, so a number is never left unattributable
- `qa_evaluator_output.json` — per-item judge scores (same shape as NAT QA output)
- `latency_summary.json` — per-item wall-clock around `vss vlm run`
- `workflow_output.json` — raw answers
- `summary.csv`

## Rules

- Drive the VLM only through `vss vlm run`. Never `POST /generate` or hand-built
  `/v1/chat/completions`.
- Do not wrap `vss` in retries. `--timeout` is the bound; the script adds only a
  hard kill 60 s past it, so a CLI that never returns cannot cost the whole run.
  A killed item is recorded as an error naming the watchdog, never as a low score.
- Items must declare `evaluation_method` containing `qa` and carry a text
  `ground_truth`. Report, trajectory-only, and unmarked items are skipped.

## Failures

Branch on the exit code; never scrape stdout for the word "error".

| Exit | Meaning | What to do |
|---|---|---|
| 0 | Every item answered | Read `summary.json` |
| 2 | Precondition wrong — a dataset flag missing, no DSS credential, no judge configured, dataset or videos not found, no QA items | Fix the setup. Re-running unchanged fails identically |
| 3 | The download failed, or at least one item errored | Read each item's `error` in `summary.json` |

A `vss` call that exits 4 (service missing from the recorded config) surfaces as an
item error, so the run ends at exit 3 — the fix is `vss configure`, not a flag.

Failures worth recognising by their message:

- `HTTP 422 ... content ... valid string` on the big clips — the recorded origin is
  loopback, so clips are being inlined as base64. Reconfigure with a routable address.
- `Did not find tenant_id` — SSO signed you in but selected no tenant. Export
  `NVDATASET_TENANTID`, or `nvdataset auth context use`.
- `LLM judge HTTP 403 ... key_model_access_denied` or `400 Invalid model name` on
  every item — the judge id is not what that gateway calls the model. Gateways that
  front several providers usually want a fully-qualified id and reject the bare name.
  `GET <judge-base-url>/models` lists the ids the key may use; copy one verbatim into
  `EVAL_LLM_JUDGE_NAME`. The VLM answers are unaffected, so only scoring is lost.
- An item error naming the watchdog — the CLI never returned and was killed at
  `--timeout` + 60 s. That is recorded as an error, never as a low score. Do not retry.
- Accuracy far from the ~0.465 baseline is not a harness failure. The judge model and
  `--num-frames` both move it; check `judge_model` and `model_served` before filing.

Implementation: [`scripts/benchmark_vlm_qa.py`](scripts/benchmark_vlm_qa.py), tested by
[`scripts/tests/`](scripts/tests/).
Dataset download contract: [`README_eval.md`](../../../deploy/docker/developer-profiles/dev-profile-base/eval/README_eval.md).