Treating a code as a replacement verdict
A support report should document observed faults and controls, not promise an RMA or refund.
Operator task guide
An Xid identifies an event family, not a complete root cause or permission to reset hardware. Preserve the full line, GPU identity, preceding application error and peer events before choosing recovery.
Start with the affected operation and the earliest event. A peer NVLink error can follow another GPU failing, while a software channel error can leave the GPU healthy. Use the driver and GPU generation appropriate recovery workflow, then rerun the workload.
Reviewed . Reference guidance is not a diagnosis of your workload.
Collect the complete message including subcode and payload, timestamp, GPU UUID, PCI address, driver version and workload traceback. If permissions prevent host collection, give this exact list and time window to the administrator.
nvidia-smi --query-gpu=index,uuid,pci.bus_id,driver_version --format=csv
# Host kernel access may require an administrator.
journalctl -k --since '30 minutes ago' | grep -iE 'NVRM|Xid|PCIe|AER'For 43 inspect the application fault. For 62 preserve micro-controller and driver context. For 79 compare PCIe and system events. For 74 or 149 inspect both link endpoints; 149 needs its NETIR subcode, not only the number. Correlate clocks before assuming simultaneous faults share an initiator.
nvidia-smi topo -m
nvidia-smi nvlink -s
# Compare with expected topology. An inactive link alone is not a fault.Follow the platform and NVIDIA recovery action for that event and release. Device resets, node restarts and power cycles interrupt workloads. Confirm device identity, drain decisions, checkpoint state and other users with the operator first. A diagnostic guide does not authorize those actions.
After the approved action, check enumeration and the affected link where relevant. Rerun the operation that failed, compare expected results, check training progress and finite gradients when applicable, and watch for the event returning. A clean device listing alone is not recovery proof.
| Signal | What it means | Next action |
|---|---|---|
| 43: GPU stopped processing | Often an application-induced channel fault; the GPU can remain healthy. | Find the first application or kernel exception and reproduce the operation before requesting replacement. |
| 62: micro-controller halt | Controller state needs the release-specific recovery workflow. | Preserve the driver and preceding events; investigate recurrence after approved recovery. |
| 74: NVLink error | The link or its remote endpoint may be affected. | Compare peer events and link identity before assigning the defective component. |
| 79: GPU inaccessible over PCIe | Reachability is lost; link, GPU and driver evidence can differ. | Collect PCIe/AER and system events, then follow the platform restart procedure. |
| 149: NVLink NETIR event | The subcode and payload determine the investigation branch. | Preserve the complete NETIR line and peer evidence. Do not reuse a generic Xid 74 remedy blindly. |
A support report should document observed faults and controls, not promise an RMA or refund.
Transient evidence can disappear. Coordinate collection and recovery with the operator.
Find the exact event family and the linked incident entry.
Use recurrence and ownership controls before a replacement request.
Separate survivor timeouts from earlier failures.
No. Some events concern application behavior, and link faults can be reported by a healthy peer. Collect the full event and relevant control before assigning ownership.
No. Collection is separate from recovery. An authorized operator must decide whether a disruptive action is appropriate for the affected workload and platform.
Use the three free diagnoses to review your error and relevant evidence. Keep reference guidance separate from the cause and recovery status of your own workload.
Diagnose your incident