How to identify the first failed rank
Find the first failed rank in distributed GPU logs by separating initiating CUDA, host and application evidence from NCCL timeout and launcher cascades.
Open the guideOperator guides
These guides start with the decision an operator must make, preserve the evidence that disappears during recovery, and end with a control that proves ownership.
Find the first failed rank in distributed GPU logs by separating initiating CUDA, host and application evidence from NCCL timeout and launcher cascades.
Open the guideDecide whether a GPU needs RMA using Xid, ECC, UUID, recurrence and controlled card, slot, node and workload evidence without replacing a healthy board.
Open the guideDistinguish Slurm cgroup host-memory OOM, CUDA GPU-memory OOM, step-creation limits and secondary NCCL timeouts using sacct, scontrol and rank evidence.
Open the guide