get-available-resources · diff

git:20260611.1b8fae3 to v1.1

192 added, 206 removed. Audit A to A.

---
name: get-available-resources
- description: This skill should be used at the start of any computationally intensive scientific task to detect and report available system resources (CPU cores, GPUs, memory, disk space). It creates a JSON file with resource information and strategic recommendations that inform computational approach decisions such as whether to use parallel processing (joblib, multiprocessing), out-of-core computing (Dask, Zarr), GPU acceleration (PyTorch, JAX), or memory-efficient strategies. Use this skill before running analyses, training models, processing large datasets, or any task where resource constraints matter.
- license: MIT license
- metadata: {"version": "1.0", "skill-author": "K-Dense Inc."}
+ description: Detect host inventory and effective CPU, memory, disk, scheduler, container, and accelerator limits when a user asks for resource-aware planning or before a clearly resource-sensitive local workload. Produces a redacted JSON snapshot and conservative planning helpers without stress tests or assuming visible host hardware is usable.
+ license: MIT
+ compatibility: Python 3.11+ on Linux, macOS, or Windows; standard library by default, optional psutil 7.2.2; accelerator and scheduler CLIs are optional read-only probes.
+ metadata:
+ version: "1.1"
+ skill-author: K-Dense Inc.
---
# Get Available Resources
- ## Overview
+ Build a conservative picture of resources available to the **current process**.
+ Keep host inventory, process affinity, cgroup/container limits, scheduler
+ allocation, and accelerator runtime usability separate.
- Detect available computational resources and generate strategic recommendations for scientific computing tasks. This skill automatically identifies CPU capabilities, GPU availability (NVIDIA CUDA, AMD ROCm, Apple Silicon Metal), memory constraints, and disk space to help make informed decisions about computational approaches.
+ ## Safety contract
- ## When to Use This Skill
+ Follow these rules:
- Use this skill proactively before any computationally intensive task:
+ - Run detection when the user requests it or a specific workload needs resource
+ planning. Do not persist a fingerprint for every scientific task.
+ - Use stdout by default. Persist only when the user chooses an explicit generic
+ local filename.
+ - Do not run stress tests, benchmarks, large allocations, write probes, device
+ resets, driver installation, or clock/power changes.
+ - Do not dump the environment. Read only the named Slurm and accelerator
+ variables implemented by the detector.
+ - Do not report hostnames, absolute paths, cgroup paths, job IDs, device UUIDs,
+ PCI addresses, or raw visibility-variable values.
+ - Treat a missing observation as unknown. Never convert unknown to unlimited.
+ - Never infer that a visible host CPU, memory pool, or GPU is usable inside a
+ scheduler allocation or container.
- - **Before data analysis**: Determine if datasets can be loaded into memory or require out-of-core processing
- - **Before model training**: Check if GPU acceleration is available and which backend to use
- - **Before parallel processing**: Identify optimal number of workers for joblib, multiprocessing, or Dask
- - **Before large file operations**: Verify sufficient disk space and appropriate storage strategies
- - **At project initialization**: Understand baseline capabilities for making architectural decisions
+ The bundled detector uses only fixed executable/argument tuples, no shell,
+ short timeouts, bounded stdout/stderr, and partial-failure warnings.
- **Example scenarios:**
- - "Help me analyze this 50GB genomics dataset" → Use this skill first to determine if Dask/Zarr are needed
- - "Train a neural network on this data" → Use this skill to detect available GPUs and backends
- - "Process 10,000 files in parallel" → Use this skill to determine optimal worker count
- - "Run a computationally intensive simulation" → Use this skill to understand resource constraints
+ ## Quick start
- ## How This Skill Works
+ Run from this skill directory.
- ### Resource Detection
+ ### Ephemeral stdout snapshot
- The skill runs `scripts/detect_resources.py` to automatically detect:
+ ```bash
+ python scripts/detect_resources.py
+ ```
- 1. **CPU Information**
- - Physical and logical core counts
- - Processor architecture and model
- - CPU frequency information
+ The command emits only JSON to stdout. Redirect it only when ordinary shell
+ permissions are acceptable.
- 2. **GPU Information**
- - NVIDIA GPUs: Detects via nvidia-smi, reports VRAM, driver version, compute capability
- - AMD GPUs: Detects via rocm-smi
- - Apple Silicon: Detects M1/M2/M3/M4 chips with Metal support and unified memory
+ ### Explicit private file
- 3. **Memory Information**
- - Total and available RAM
- - Current memory usage percentage
- - Swap space availability
+ ```bash
+ python scripts/detect_resources.py --output resource-snapshot.json
+ ```
- 4. **Disk Space Information**
- - Total and available disk space for working directory
- - Current usage percentage
+ Explicit output is restricted to one `.json` filename in the current
+ directory, uses private permissions, rejects symlinks and path traversal, and
+ refuses overwrite unless `--force` is supplied.
- 5. **Operating System Information**
- - OS type (macOS, Linux, Windows)
- - OS version and release
- - Python version
+ ### Optional psutil enhancement
- ### Output Format
+ The standard-library detector works without installation. For broader
+ cross-platform physical-core, affinity, available-memory, swap, and disk
+ coverage:
- The skill generates a `.claude_resources.json` file in the current working directory containing:
+ ```bash
+ uv pip install "psutil==7.2.2"
+ ```
- ```json
- {
- "timestamp": "2025-10-23T10:30:00",
- "os": {
- "system": "Darwin",
- "release": "25.0.0",
- "machine": "arm64"
- },
- "cpu": {
- "physical_cores": 8,
- "logical_cores": 8,
- "architecture": "arm64"
- },
- "memory": {
- "total_gb": 16.0,
- "available_gb": 8.5,
- "percent_used": 46.9
- },
- "disk": {
- "total_gb": 500.0,
- "available_gb": 200.0,
- "percent_used": 60.0
- },
- "gpu": {
- "nvidia_gpus": [],
- "amd_gpus": [],
- "apple_silicon": {
- "name": "Apple M2",
- "type": "Apple Silicon",
- "backend": "Metal",
- "unified_memory": true
- },
- "total_gpus": 1,
- "available_backends": ["Metal"]
- },
- "recommendations": {
- "parallel_processing": {
- "strategy": "high_parallelism",
- "suggested_workers": 6,
- "libraries": ["joblib", "multiprocessing", "dask"]
- },
- "memory_strategy": {
- "strategy": "moderate_memory",
- "libraries": ["dask", "zarr"],
- "note": "Consider chunking for datasets > 2GB"
- },
- "gpu_acceleration": {
- "available": true,
- "backends": ["Metal"],
- "suggested_libraries": ["pytorch-mps", "tensorflow-metal", "jax-metal"]
- },
- "large_data_handling": {
- "strategy": "disk_abundant",
- "note": "Sufficient space for large intermediate files"
- }
- }
- }
+ The import is lazy. Failure to import psutil becomes a warning, not a fatal
+ error.
+
+ ### Skip management-tool probes
+
+ ```bash
+ python scripts/detect_resources.py --skip-accelerators
```
- ### Strategic Recommendations
+ Use this when accelerator discovery latency is undesirable. The detector still
+ summarizes the presence and state of allowlisted visibility variables without
+ returning their values.
- The skill generates context-aware recommendations:
+ ## Required interpretation
- **Parallel Processing Recommendations:**
- - **High parallelism (8+ cores)**: Use Dask, joblib, or multiprocessing with workers = cores - 2
- - **Moderate parallelism (4-7 cores)**: Use joblib or multiprocessing with workers = cores - 1
- - **Sequential (< 4 cores)**: Prefer sequential processing to avoid overhead
+ ### CPU
- **Memory Strategy Recommendations:**
- - **Memory constrained (< 4GB available)**: Use Zarr, Dask, or H5py for out-of-core processing
- - **Moderate memory (4-16GB available)**: Use Dask/Zarr for datasets > 2GB
- - **Memory abundant (> 16GB available)**: Can load most datasets into memory directly
+ Read these as different facts:
- **GPU Acceleration Recommendations:**
- - **NVIDIA GPUs detected**: Use PyTorch, TensorFlow, JAX, CuPy, or RAPIDS
- - **AMD GPUs detected**: Use PyTorch-ROCm or TensorFlow-ROCm
- - **Apple Silicon detected**: Use PyTorch with MPS backend, TensorFlow-Metal, or JAX-Metal
- - **No GPU detected**: Use CPU-optimized libraries
+ - `cpu.host.logical`: system-visible scheduling units.
+ - `cpu.host.physical`: physical topology, or null; never inferred from logical
+ count.
+ - `cpu.process.affinity_logical`: current affinity-set size when supported.
+ - `cpu.cgroup_v2.cpuset_logical`: effective cgroup cpuset size.
+ - `cpu.cgroup_v2.quota_cores`: finite `cpu.max` capacity, possibly fractional.
+ - `scheduler.allocation.cpu_per_process`: bounded Slurm per-task
+ interpretation when scope is clear.
+ - `cpu.effective.capacity_cores`: minimum positive observed constraint.
+ - `cpu.effective.worker_ceiling`: conservative floor for CPU process workers.
- **Large Data Handling Recommendations:**
- - **Disk constrained (< 10GB)**: Use streaming or compression strategies
- - **Moderate disk (10-100GB)**: Use Zarr, H5py, or Parquet formats
- - **Disk abundant (> 100GB)**: Can create large intermediate files freely
+ A quota of 1.5 is CPU-time capacity, not 1.5 physical cores. Affinity and
+ cpusets constrain placement; quota constrains bandwidth.
- ## Usage Instructions
+ ### Memory
- ### Step 1: Run Resource Detection
+ Keep these separate:
- Execute the detection script at the start of any computationally intensive task:
+ - host total/available memory;
+ - current cgroup usage, hard `memory.max`, and remaining hierarchical capacity;
+ - `memory.high`, which is a pressure/throttle boundary rather than a hard cap;
+ - scheduler memory allocation and its scope; and
+ - conservative effective hard limit and available estimate.
- ```bash
- python scripts/detect_resources.py
- ```
+ On Apple silicon, `memory.model` is `unified_cpu_gpu`. Do not add integrated GPU
+ memory to RAM or describe it as separate VRAM.
- Optional arguments:
- - `-o, --output <path>`: Specify custom output path (default: `.claude_resources.json`)
- - `-v, --verbose`: Print full resource information to stdout
+ ### Accelerators
- ### Step 2: Read and Apply Recommendations
+ Each device is a backend **candidate**:
- After running detection, read the generated `.claude_resources.json` file to inform computational decisions:
+ - NVIDIA GPU → CUDA candidate;
+ - AMD GPU → ROCm candidate;
+ - Apple integrated GPU → Metal candidate.
- ```python
- # Example: Use recommendations in code
- import json
+ Management-query visibility does not establish:
- with open('.claude_resources.json', 'r') as f:
- resources = json.load(f)
+ 1. scheduler/container permission;
+ 2. device-node access;
+ 3. driver/runtime compatibility;
+ 4. framework package compatibility; or
+ 5. operator/data-type support.
- # Check parallel processing strategy
- if resources['recommendations']['parallel_processing']['strategy'] == 'high_parallelism':
- n_jobs = resources['recommendations']['parallel_processing']['suggested_workers']
- # Use joblib, Dask, or multiprocessing with n_jobs workers
+ Therefore `runtime_usable_devices` remains null and each device says
+ `runtime_compatibility: not_tested`. Visibility/allocation counts are upper
+ bounds, not guarantees.
- # Check memory strategy
- if resources['recommendations']['memory_strategy']['strategy'] == 'memory_constrained':
- # Use Dask, Zarr, or H5py for out-of-core processing
- import dask.array as da
- # Load data in chunks
+ ### Disk
- # Check GPU availability
- if resources['recommendations']['gpu_acceleration']['available']:
- backends = resources['recommendations']['gpu_acceleration']['backends']
- # Use appropriate GPU library based on available backend
- ```
+ `capacity_bytes`, filesystem `free_bytes`, user-available blocks, and a
+ non-writing permission check are distinct. Filesystem or project quotas can
+ still be stricter. The absolute working path is always redacted.
- ### Step 3: Make Informed Decisions
+ ### Scheduler and container
- Use the resource information and recommendations to make strategic choices:
+ Slurm variables describe allocation scope, but enforcement depends on site
+ configuration such as task affinity or cgroups. Prefer affinity and cgroup
+ observations as enforcement evidence.
- **For data loading:**
- ```python
- memory_available_gb = resources['memory']['available_gb']
- dataset_size_gb = 10
+ Container markers identify context; cgroup controls identify limits. A
+ container with no finite cgroup value can still see host inventory, and a
+ non-root cgroup is not automatically labeled a container.
- if dataset_size_gb > memory_available_gb * 0.5:
- # Dataset is large relative to memory, use Dask
- import dask.dataframe as dd
- df = dd.read_csv('large_file.csv')
- else:
- # Dataset fits in memory, use pandas
- import pandas as pd
- df = pd.read_csv('large_file.csv')
- ```
+ See [`references/resource_semantics.md`](references/resource_semantics.md) for
+ the detailed platform rules.
- **For parallel processing:**
- ```python
- from joblib import Parallel, delayed
+ ## Plan a workload
- n_jobs = resources['recommendations']['parallel_processing'].get('suggested_workers', 1)
+ The planner consumes a validated snapshot and performs no work:
- results = Parallel(n_jobs=n_jobs)(
- delayed(process_function)(item) for item in data
- )
+ ```bash
+ python scripts/plan_workload.py resource-snapshot.json \
+ --workload cpu \
+ --tasks 100 \
+ --memory-per-worker-mib 2048
```
- **For GPU acceleration:**
- ```python
- import torch
+ Optional controls:
- if 'CUDA' in resources['gpu']['available_backends']:
- device = torch.device('cuda')
- elif 'Metal' in resources['gpu']['available_backends']:
- device = torch.device('mps')
- else:
- device = torch.device('cpu')
+ - `--workers N`: explicit upper bound.
+ - `--reserve-memory-mib N`: memory kept outside the worker budget.
+ - `--workload cpu|mixed|io`: selects a bounded worker heuristic.
+ - `--accelerator none|any|cuda|rocm|metal`: requests a candidate backend
+ decision without claiming usability.
+ - `--output plan.json`: explicit private local output; stdout is default.
- model = model.to(device)
+ For CPU or mixed work, use `suggested_workers` and
+ `threads_per_worker` together. Process workers multiplied by BLAS/OpenMP native
+ threads can oversubscribe an allocation.
+
+ The I/O plan permits bounded oversubscription (maximum 32) but labels it a
+ heuristic. Benchmark only the real representative workload and stay within
+ scheduler/container limits.
+
+ ## Validate or diff snapshots
+
+ Validate:
+
+ ```bash
+ python scripts/snapshot_tools.py validate resource-snapshot.json
```
- ## Dependencies
+ Diff resource state while ignoring `observed_at`:
- The detection script requires the following Python packages:
+ ```bash
+ python scripts/snapshot_tools.py diff before.json after.json
+ ```
+ Use `--include-volatile` to include the timestamp. Inputs must be regular,
+ non-symlink JSON files no larger than 1 MiB. Diffs are bounded.
+
+ The schema and null/zero meanings are documented in
+ [`references/snapshot_schema.md`](references/snapshot_schema.md).
+
+ ## Optional accelerator diagnostic plan
+
+ Generate a plan without executing any diagnostic:
+
```bash
- uv pip install psutil
+ python scripts/accelerator_diagnostics.py resource-snapshot.json \
+ --backend auto
```
- All other functionality uses Python standard library modules (json, os, platform, subprocess, sys, pathlib).
+ The result contains fixed, read-only management query argument lists and
+ separate gates for visibility, permission, and runtime compatibility. Run a
+ framework's official availability check only in the exact environment that
+ will execute the workload. Do not install or mutate drivers automatically.
- ## Platform Support
+ ## Partial failures and provenance
- - **macOS**: Full support including Apple Silicon (M1/M2/M3/M4) GPU detection
- - **Linux**: Full support including NVIDIA (nvidia-smi) and AMD (rocm-smi) GPU detection
- - **Windows**: Full support including NVIDIA GPU detection
+ One failed probe must not erase successful observations. Inspect:
- ## Best Practices
+ - `completeness`;
+ - sorted `warnings` with stable codes;
+ - sorted `provenance` source/status records; and
+ - null fields.
- 1. **Run early**: Execute resource detection at the start of projects or before major computational tasks
- 2. **Re-run periodically**: System resources change over time (memory usage, disk space)
- 3. **Check before scaling**: Verify resources before scaling up parallel workers or data sizes
- 4. **Document decisions**: Keep the `.claude_resources.json` file in project directories to document resource-aware decisions
- 5. **Use with versioning**: Different machines have different capabilities; resource files help maintain portability
+ Subprocess stderr and raw exception text are not copied into the snapshot
+ because they can contain identifiers or paths.
- ## Troubleshooting
+ ## Platform notes
- **GPU not detected:**
- - Ensure GPU drivers are installed (nvidia-smi, rocm-smi, or system_profiler for Apple Silicon)
- - Check that GPU utilities are in system PATH
- - Verify GPU is not in use by other processes
+ - **Linux:** reads only bounded `/proc` and cgroup v2 files. Ancestor CPU and
+ memory limits are considered.
+ - **macOS:** uses fixed `sysctl` keys and a bounded
+ `system_profiler SPDisplaysDataType -json` query. Apple silicon memory is
+ unified.
+ - **Windows:** optional psutil improves physical-core, affinity, available
+ memory, and swap observations. Processor-group scope can make host and
+ process counts differ.
+ - **Slurm:** reads an allowlist of allocation variables. It never emits job,
+ node, submit-host, GPU-ID, or path values.
+ - **NVIDIA/AMD:** management CLIs are optional. Absence is normal; timeout,
+ truncation, parse failure, and runtime uncertainty remain explicit.
- **Script execution fails:**
- - Ensure psutil is installed: `uv pip install psutil`
- - Check Python version compatibility (Python 3.6+)
- - Verify script has execute permissions: `chmod +x scripts/detect_resources.py`
+ ## Bundled files
- **Inaccurate memory readings:**
- - Memory readings are snapshots; actual available memory changes constantly
- - Close other applications before detection for accurate "available" memory
- - Consider running detection multiple times and averaging results
+ - `scripts/detect_resources.py` — redacted snapshot collector.
+ - `scripts/plan_workload.py` — deterministic worker/memory planner.
+ - `scripts/snapshot_tools.py` — schema validator and bounded structural diff.
+ - `scripts/accelerator_diagnostics.py` — non-executing read-only diagnostic
+ plan.
+ - `tests/test_scripts.py` and `tests/fixtures/resource_cases.json` —
+ network-free Linux, macOS, Windows, cgroup, Slurm, and accelerator cases.
+ - `references/resource_semantics.md` — interpretation and platform details.
+ - `references/snapshot_schema.md` — schema 1.1 contract.
+ - `references/sources.md` — dated official-source ledger.
+ Official documentation was refreshed on **2026-07-23**; consult
+ [`references/sources.md`](references/sources.md) before changing semantics or
+ dependency pins.