Free tools for GPU cluster engineers
Diagnostic references, local parsers and planning calculators that run without an account. Each tool states what it can establish and where a complete incident still needs more evidence.
- Hardware, software or more evidence
NVIDIA Xid Decoder
Paste an NVRM Xid line to extract the code, fault class, RMA relevance and first evidence-preserving action, then open the complete investigation.
Open NVIDIA Xid Decoder - Turn a GPU report into incident evidence
nvidia-smi Parser
Parse GPU identity, memory, thermal, power, utilization, ECC and row-remapper fields locally in your browser. Flags stay evidence, not automatic hardware verdicts.
Open nvidia-smi Parser - Find the first evidence, not the last timeout
NCCL Log Decoder
Classify covered CUDA, bootstrap, transport, duplicate-GPU and timeout signatures across rank logs, with wrappers marked as secondary.
Open NCCL Log Decoder - Model direct capacity and engineer time
GPU Downtime Calculator
Calculate GPU hours, infrastructure cost, labor, monthly exposure and a scenario-based recoverable share using your own inputs.
Open GPU Downtime Calculator - Catch fabric misconfigs before step 0
NCCL Topology Linter
Paste nvidia-smi topo -m and your NCCL_* env. It checks GPU→NIC affinity, PCIe ACS, memlock limits and NCCL settings against the fabric you actually have, and names the ones that will cost you throughput or hang the job.
Open NCCL Topology Linter - Find what changed since the last good run
Cross-Run Environment Diff
Paste artifacts from the last good run and the failing one. Driver, CUDA, NCCL, PyTorch/JAX, pip packages and NCCL_* env are diffed, and every change is ranked against the observed failure class so you know where to bisect first.
Open Cross-Run Environment Diff - Stop losing hours to env vars
NCCL Environment Generator
Describe your interconnect and job shape and get a defensible NCCL and PyTorch environment, with each variable explained so you can justify it in review instead of copying a Stack Overflow block.
Open NCCL Environment Generator - The failure data procurement never gets
GPU Reliability Index
H100 vs A100 vs H200 failure modes and rates drawn from the Denpex incident corpus, so a hardware decision can cite something other than a vendor datasheet.
Open GPU Reliability Index