Skip to content

Operator guides

GPU incident guides and controlled recovery case studies

Find the evidence that distinguishes possible causes and the checks needed after a recommendation. The case studies retain unsuccessful instructions and partial repairs alongside bounded recovery results.

Diagnostic method

How to identify the first failed rank

Find the first failed rank in distributed GPU logs by separating initiating CUDA, host and application evidence from NCCL timeout and launcher cascades.

Open the guide

Diagnostic method

How to determine whether a GPU needs RMA

Decide whether a GPU needs RMA using Xid, ECC, UUID, recurrence and controlled card, slot, node and workload evidence without replacing a healthy board.

Open the guide

Diagnostic method

How to diagnose a Slurm GPU OOM

Distinguish Slurm cgroup host-memory OOM, CUDA GPU-memory OOM, step-creation limits and secondary NCCL timeouts using sacct, scontrol and rank evidence.

Open the guide

Diagnostic method

How to diagnose silent GPU stragglers

Diagnose silent GPU stragglers and barrier slowdowns by isolating thermal throttling, PCIe bandwidth drops, and rank execution timing in distributed training.

Open the guide

Diagnostic method

How to diagnose multi-node cluster drift

Diagnose multi-node cluster configuration drift by comparing PCIe link widths, NVIDIA driver versions, NUMA bindings, and network settings across peer hosts.

Open the guide

Diagnostic method

Which evidence distinguishes Xid 43, 62, 74, 79 and 149?

Distinguish application, GPU reachability and NVLink evidence before restarting a job, resetting a device or escalating an Xid incident.

Open the guide

Diagnostic method

Fix the device mismatch at the failing PyTorch operation

Check input, parameters, buffers, indices and restored optimizer state at the failing PyTorch operation, then verify the intended output and update.

Open the guide

Diagnostic method

Use Compute Sanitizer on the smallest failing CUDA workload

Use a reduced CUDA workload and memcheck to locate memory errors, retain source attribution and verify outputs after the change.

Open the guide

Diagnostic method

Diagnose safetensors HeaderTooLarge as a file-format problem

Inspect the safetensors length prefix, file identity and download integrity without parsing an unbounded header or changing GPU settings.

Open the guide

Diagnostic method

Which vLLM memory or context limit actually failed?

Separate model context limits, KV cache capacity and allocation failures using startup evidence, resolved settings and a representative request.

Open the guide

Controlled replay case study

A correct tensor edit, an incorrect launch flag and a verified replay

A controlled RTX 4090 replay preserved tensor order after a view failure and exposed an unsupported launch flag in the initial recommendation.

Read the replay record

Controlled replay case study

An optimizer fix that stopped crashing but did not yet train correctly

A controlled RTX 4090 replay found that reordering optimizer updates removed an exception but required gradient isolation to restore intended training.

Read the replay record