Diagnostic method
How to identify the first failed rank
Find the first failed rank in distributed GPU logs by separating initiating CUDA, host and application evidence from NCCL timeout and launcher cascades.
Open the guideOperator guides
Find the evidence that distinguishes possible causes and the checks needed after a recommendation. The case studies retain unsuccessful instructions and partial repairs alongside bounded recovery results.
Diagnostic method
Find the first failed rank in distributed GPU logs by separating initiating CUDA, host and application evidence from NCCL timeout and launcher cascades.
Open the guideDiagnostic method
Decide whether a GPU needs RMA using Xid, ECC, UUID, recurrence and controlled card, slot, node and workload evidence without replacing a healthy board.
Open the guideDiagnostic method
Distinguish Slurm cgroup host-memory OOM, CUDA GPU-memory OOM, step-creation limits and secondary NCCL timeouts using sacct, scontrol and rank evidence.
Open the guideDiagnostic method
Diagnose silent GPU stragglers and barrier slowdowns by isolating thermal throttling, PCIe bandwidth drops, and rank execution timing in distributed training.
Open the guideDiagnostic method
Diagnose multi-node cluster configuration drift by comparing PCIe link widths, NVIDIA driver versions, NUMA bindings, and network settings across peer hosts.
Open the guideDiagnostic method
Distinguish application, GPU reachability and NVLink evidence before restarting a job, resetting a device or escalating an Xid incident.
Open the guideDiagnostic method
Check input, parameters, buffers, indices and restored optimizer state at the failing PyTorch operation, then verify the intended output and update.
Open the guideDiagnostic method
Use a reduced CUDA workload and memcheck to locate memory errors, retain source attribution and verify outputs after the change.
Open the guideDiagnostic method
Inspect the safetensors length prefix, file identity and download integrity without parsing an unbounded header or changing GPU settings.
Open the guideDiagnostic method
Separate model context limits, KV cache capacity and allocation failures using startup evidence, resolved settings and a representative request.
Open the guideControlled replay case study
A controlled RTX 4090 replay preserved tensor order after a view failure and exposed an unsupported launch flag in the initial recommendation.
Read the replay recordControlled replay case study
A controlled RTX 4090 replay found that reordering optimizer updates removed an exception but required gradient isolation to restore intended training.
Read the replay record