Skip to content

Operator guides

High-stakes GPU incident tasks, one verified step at a time

These guides start with the decision an operator must make, preserve the evidence that disappears during recovery, and end with a control that proves ownership.

How to identify the first failed rank

Find the first failed rank in distributed GPU logs by separating initiating CUDA, host and application evidence from NCCL timeout and launcher cascades.

Open the guide

How to determine whether a GPU needs RMA

Decide whether a GPU needs RMA using Xid, ECC, UUID, recurrence and controlled card, slot, node and workload evidence without replacing a healthy board.

Open the guide

How to diagnose a Slurm GPU OOM

Distinguish Slurm cgroup host-memory OOM, CUDA GPU-memory OOM, step-creation limits and secondary NCCL timeouts using sacct, scontrol and rank evidence.

Open the guide