Free tools for GPU cluster engineers
Four tools that run without an account. They use the same failure corpus as the paid diagnosis engine, so what they tell you is what Denpex would tell you.
- Catch fabric misconfigs before step 0
NCCL Topology Linter
Paste nvidia-smi topo -m and your NCCL_* env. It checks GPU→NIC affinity, PCIe ACS, memlock limits and NCCL settings against the fabric you actually have, and names the ones that will cost you throughput or hang the job.
Open NCCL Topology Linter - Find what changed since the last good run
Cross-Run Environment Diff
Paste artifacts from the last good run and the failing one. Driver, CUDA, NCCL, PyTorch/JAX, pip packages and NCCL_* env are diffed, and every change is ranked against the observed failure class so you know where to bisect first.
Open Cross-Run Environment Diff - Stop losing hours to env vars
NCCL Environment Generator
Describe your interconnect and job shape and get a defensible NCCL and PyTorch environment, with each variable explained so you can justify it in review instead of copying a Stack Overflow block.
Open NCCL Environment Generator - The failure data procurement never gets
GPU Reliability Index
H100 vs A100 vs H200 failure modes and rates drawn from the Denpex incident corpus, so a hardware decision can cite something other than a vendor datasheet.
Open GPU Reliability Index