Skip to content

Operator task guide

Which evidence distinguishes Xid 43, 62, 74, 79 and 149?

An Xid identifies an event family, not a complete root cause or permission to reset hardware. Preserve the full line, GPU identity, preceding application error and peer events before choosing recovery.

Start with the affected operation and the earliest event. A peer NVLink error can follow another GPU failing, while a software channel error can leave the GPU healthy. Use the driver and GPU generation appropriate recovery workflow, then rerun the workload.

Reviewed . Reference guidance is not a diagnosis of your workload.

Step-by-step method

  1. 1

    Preserve the device and event sequence

    Collect the complete message including subcode and payload, timestamp, GPU UUID, PCI address, driver version and workload traceback. If permissions prevent host collection, give this exact list and time window to the administrator.

    nvidia-smi --query-gpu=index,uuid,pci.bus_id,driver_version --format=csv
    # Host kernel access may require an administrator.
    journalctl -k --since '30 minutes ago' | grep -iE 'NVRM|Xid|PCIe|AER'
  2. 2

    Choose the relevant evidence branch

    For 43 inspect the application fault. For 62 preserve micro-controller and driver context. For 79 compare PCIe and system events. For 74 or 149 inspect both link endpoints; 149 needs its NETIR subcode, not only the number. Correlate clocks before assuming simultaneous faults share an initiator.

    nvidia-smi topo -m
    nvidia-smi nvlink -s
    # Compare with expected topology. An inactive link alone is not a fault.
  3. 3

    Prepare an operator-controlled recovery

    Follow the platform and NVIDIA recovery action for that event and release. Device resets, node restarts and power cycles interrupt workloads. Confirm device identity, drain decisions, checkpoint state and other users with the operator first. A diagnostic guide does not authorize those actions.

  4. 4

    Verify the affected workload, not only nvidia-smi

    After the approved action, check enumeration and the affected link where relevant. Rerun the operation that failed, compare expected results, check training progress and finite gradients when applicable, and watch for the event returning. A clean device listing alone is not recovery proof.

Event family and the next discriminating fact

Signals, meanings and actions for which evidence distinguishes xid 43, 62, 74, 79 and 149?.
SignalWhat it meansNext action
43: GPU stopped processingOften an application-induced channel fault; the GPU can remain healthy.Find the first application or kernel exception and reproduce the operation before requesting replacement.
62: micro-controller haltController state needs the release-specific recovery workflow.Preserve the driver and preceding events; investigate recurrence after approved recovery.
74: NVLink errorThe link or its remote endpoint may be affected.Compare peer events and link identity before assigning the defective component.
79: GPU inaccessible over PCIeReachability is lost; link, GPU and driver evidence can differ.Collect PCIe/AER and system events, then follow the platform restart procedure.
149: NVLink NETIR eventThe subcode and payload determine the investigation branch.Preserve the complete NETIR line and peer evidence. Do not reuse a generic Xid 74 remedy blindly.

Evidence checklist

  • Complete Xid lines, including NETIR subcodes
  • GPU UUID and PCI address mapped to the workload
  • Driver, GPU generation and relevant firmware versions
  • Application and peer events before the first timeout
  • Approved recovery action and subsequent recurrence window

Common mistakes

Treating a code as a replacement verdict

A support report should document observed faults and controls, not promise an RMA or refund.

Resetting before capture

Transient evidence can disappear. Coordinate collection and recovery with the operator.

Frequently asked questions

Can I conclude hardware failure from one Xid?

No. Some events concern application behavior, and link faults can be reported by a healthy peer. Collect the full event and relevant control before assigning ownership.

Does this guide run resets remotely?

No. Collection is separate from recovery. An authorized operator must decide whether a disruptive action is appropriate for the affected workload and platform.

Apply the method to your incident

Use the three free diagnoses to review your error and relevant evidence. Keep reference guidance separate from the cause and recovery status of your own workload.

Diagnose your incident