False hardware attribution
Prevent application, driver, power, thermal, riser and fabric faults from being presented as a defective board.
For hardware support and integration teams
A useful GPU RMA workflow must answer two questions before generating paperwork: does the evidence meet a hardware threshold, and does the fault follow the card rather than the slot, node, power path, driver or workload?
Denpex organizes Xid, ECC, NVLink, device identity, recurrence and verification evidence into a vendor-facing incident path. Ambiguous cases remain field-diagnosis candidates, while software-owned failures are explicitly kept out of the RMA queue.
These incidents look adjacent in the final alert, but they require different owners and different proof before recovery.
Prevent application, driver, power, thermal, riser and fabric faults from being presented as a defective board.
Attach GPU UUID, serial, PCI bus ID, host and firmware context before the hardware is removed.
Preserve kernel, BMC, Xid, ECC and recurrence evidence before power cycling clears the most useful state.
Package completed checks and the remaining uncertainty so the next escalation does not repeat basic collection.
Map the initiating signature to application, environment, fabric, configuration, hardware or unresolved ownership.
Record device identity, health state, recurrence and shared-node signals before hardware movement.
Test whether the failure follows the card, slot, node, workload or software image.
Include the verdict, evidence checklist, commands, timestamps and unresolved items in the vendor case.
Use your incidents and timestamps. These are measurement definitions, not Denpex customer-outcome claims.
| Metric | Measurement |
|---|---|
| False RMA avoidance | Suspected boards cleared by a controlled test that identified slot, node, power, thermal, driver or workload ownership. |
| Vendor acceptance cycle | Support exchanges required before the case contains accepted device identity, health evidence and reproduction data. |
| Time to hardware verdict | Elapsed time from first incident to RMA candidate, field diagnosis or not-RMA-worthy classification. |
| Repeat device incidents | Verified hardware-family events per GPU UUID after repair, replacement or return to service. |
See fault class, first action and RMA relevance for common driver-reported codes.
Review sourced failure surfaces across accelerators, memory and interconnects.
Separate a missing GPU from power, thermal, PCIe, slot and board ownership.
Export per-node events, MTTR and measured cost data to existing reporting.
No. It proves the driver lost communication with the device, but a power event, thermal cutoff, riser, slot, baseboard or dead card can all produce that symptom. The cheapest discriminator is whether the failure follows the card into a known-good slot.
Preserve the kernel log, NVIDIA bug report, GPU UUID and serial, PCI bus ID, ECC and row-remapping data, NVLink state, BMC power and thermal events, node topology, workload timestamp and any simultaneous failures on peer GPUs.
The fleet workflow can assemble an evidence payload and support-ready text from the available incident data. Missing fields remain marked with collection instructions, and the verdict can explicitly say field diagnosis first or not RMA-worthy.
Track actual boards withheld after a non-card cause was verified, support exchanges per accepted case, hardware verdict time and downtime by GPU UUID. Keep measured outcomes separate from modeled replacement or GPU-hour costs.
Freeze the expected owner and outcome, replay the evidence, then compare operator time and decision quality with your current process.
Build the evaluation plan