Skip to content

For hardware support and integration teams

OEM, systems integrator and GPU RMA workflows

A useful GPU RMA workflow must answer two questions before generating paperwork: does the evidence meet a hardware threshold, and does the fault follow the card rather than the slot, node, power path, driver or workload?

Denpex organizes Xid, ECC, NVLink, device identity, recurrence and verification evidence into a vendor-facing incident path. Ambiguous cases remain field-diagnosis candidates, while software-owned failures are explicitly kept out of the RMA queue.

Who this is for

  • GPU OEM support teams
  • Systems integrators
  • Hardware qualification engineers
  • Field service engineers
  • Data center operations teams
  • GPU cloud escalation teams

The failure surface

These incidents look adjacent in the final alert, but they require different owners and different proof before recovery.

False hardware attribution

Prevent application, driver, power, thermal, riser and fabric faults from being presented as a defective board.

Missing device identity

Attach GPU UUID, serial, PCI bus ID, host and firmware context before the hardware is removed.

Evidence lost during reset

Preserve kernel, BMC, Xid, ECC and recurrence evidence before power cycling clears the most useful state.

Vendor support loops

Package completed checks and the remaining uncertainty so the next escalation does not repeat basic collection.

The operating loop

  1. 1

    Classify the error

    Map the initiating signature to application, environment, fabric, configuration, hardware or unresolved ownership.

  2. 2

    Collect board and host evidence

    Record device identity, health state, recurrence and shared-node signals before hardware movement.

  3. 3

    Run the ownership control

    Test whether the failure follows the card, slot, node, workload or software image.

  4. 4

    Generate the support packet

    Include the verdict, evidence checklist, commands, timestamps and unresolved items in the vendor case.

What Denpex contributes

  • NVIDIA Xid classification with explicit hardware and non-hardware fault classes
  • Evidence-gated RMA payloads using node and GPU incident history
  • Verification commands and missing-evidence instructions before escalation
  • Fleet metrics for repeated hardware events and measured downtime

Measure this in an evaluation

Use your incidents and timestamps. These are measurement definitions, not Denpex customer-outcome claims.

Recommended evaluation metrics and how to measure them.
MetricMeasurement
False RMA avoidanceSuspected boards cleared by a controlled test that identified slot, node, power, thermal, driver or workload ownership.
Vendor acceptance cycleSupport exchanges required before the case contains accepted device identity, health evidence and reproduction data.
Time to hardware verdictElapsed time from first incident to RMA candidate, field diagnosis or not-RMA-worthy classification.
Repeat device incidentsVerified hardware-family events per GPU UUID after repair, replacement or return to service.

Frequently asked questions

Does Xid 79 automatically justify a GPU RMA?

No. It proves the driver lost communication with the device, but a power event, thermal cutoff, riser, slot, baseboard or dead card can all produce that symptom. The cheapest discriminator is whether the failure follows the card into a known-good slot.

What evidence should be collected before a GPU reset?

Preserve the kernel log, NVIDIA bug report, GPU UUID and serial, PCI bus ID, ECC and row-remapping data, NVLink state, BMC power and thermal events, node topology, workload timestamp and any simultaneous failures on peer GPUs.

Can Denpex generate a vendor support case?

The fleet workflow can assemble an evidence payload and support-ready text from the available incident data. Missing fields remain marked with collection instructions, and the verdict can explicitly say field diagnosis first or not RMA-worthy.

How do we measure whether the workflow saves money?

Track actual boards withheld after a non-card cause was verified, support exchanges per accepted case, hardware verdict time and downtime by GPU UUID. Keep measured outcomes separate from modeled replacement or GPU-hour costs.

Prove it on incidents your team already resolved

Freeze the expected owner and outcome, replay the evidence, then compare operator time and decision quality with your current process.

Build the evaluation plan