Treating Xid 79 as a dead board
The driver lost the device, but power, thermal, riser, slot, baseboard and card faults all satisfy that description.
Operator task guide
A GPU needs RMA when repeatable, device-owned hardware evidence survives reasonable software, node and slot controls. One ambiguous Xid, CUDA crash or missing device is not enough because power, cooling, PCIe, driver and application faults can create the same outer symptom.
Preserve device identity and host evidence before resetting, classify the Xid or ECC event, check recurrence on the same serial, and run the cheapest test that separates the card from its environment. Then use one of three honest outcomes: RMA candidate, field diagnosis first, or not RMA-worthy.
Kernel, BMC and driver evidence can disappear during recovery. Record the GPU UUID, serial, PCI bus ID, node, timestamp and every simultaneous event before touching the device.
dmesg -T | rg -i 'nvrm|xid|ecc|nvlink|pcie'
nvidia-bug-report.sh
nvidia-smi --query-gpu=index,uuid,serial,pci.bus_id,driver_version,vbios_version --format=csvUncorrectable ECC, failed row remapping, repeated microcontroller halts and persistent link loss carry more hardware weight than application-class Xids or a generic CUDA error. Read the event in its full sequence, not as an isolated number.
Reproduce the workload with a known-good driver and image or with a sanitizer when the Xid class points to application memory access. A failure that disappears under serialization or follows the input is not strong board evidence.
Repeated events on one serial strengthen card ownership. Simultaneous events across several GPUs weaken it and point toward shared power, cooling, baseboard, PCIe or NVSwitch infrastructure.
nvidia-smi -q -d ECC,ROW_REMAPPER,PAGE_RETIREMENT
nvidia-smi nvlink -eWhen operational policy permits, move the suspect card to a known-good slot and put a known-good card in the suspect path. A fault that follows the card supports RMA; one that remains with the slot or node does not.
State the verdict, causal evidence, controlled test, recurrence, device identity, commands run and anything still missing. Do not turn an unresolved case into a hardware claim to make the ticket look complete.
| Signal | What it means | Next action |
|---|---|---|
| Uncorrectable ECC or failed row remapping that recurs on one serial | Strong device-memory evidence | Keep drained, protect checkpoints and prepare the vendor packet. |
| Xid 79, GPU fallen off the bus | Device communication is gone, but ownership is still ambiguous | Check BMC power, thermal and PCIe evidence, then test whether the failure follows the card. |
| Xid 13, 31, 43 or device-side assert tied to one workload | Application or software evidence | Use sanitizer or request controls. Do not RMA from this signal alone. |
| Several GPUs fail at the same timestamp | Shared node or fabric cause is more likely than independent cards | Investigate power, baseboard, cooling, PCIe and NVSwitch before replacing boards. |
| Fault stays with slot after a card swap | The card is not the recurring owner | Repair the slot, riser, node or shared path and keep the healthy board out of the RMA queue. |
The driver lost the device, but power, thermal, riser, slot, baseboard and card faults all satisfy that description.
A successful reset can erase the only device and host evidence the vendor needs, without proving the cause was fixed.
Indexes change across boots, containers and scheduler steps. Correlate recurrence by UUID or serial, with PCI bus and host as location evidence.
Paste the code and review fault class, RMA relevance and first action.
Extract identity, memory, thermal, power and ECC evidence in browser.
Follow the complete GPU-fallen-off-bus decision path.
Connect field diagnosis, vendor evidence and outcome measurement.
No. Xid 79 means the driver can no longer communicate with the GPU. A dead board is one cause, but power, thermal, riser, slot, baseboard and PCIe failures can look identical. Preserve evidence and test whether the fault follows the card.
Recurring uncorrectable memory errors, exhausted or failed row remapping, persistent microcontroller failures, and link or initialization faults that follow one serial after node and software controls provide the strongest operational evidence.
Yes. nvidia-smi proves the driver can query the device at that moment. Intermittent memory, PCIe, NVLink, power or thermal failures may disappear at idle. Use recurrence plus a production-shaped health control and preserve the prior fault evidence.
The OEM or vendor makes the warranty decision under its support terms. Denpex can classify and package operational evidence, mark missing fields and screen out not-RMA-worthy cases, but it does not certify a board as defective.
Paste the logs for an answer-first diagnosis, evidence request and verification step.
Diagnose the incident