Skip to content

Xid 79: GPU has fallen off the bus

Xid 79 means the NVIDIA driver cannot reach the GPU over PCIe. Investigate link, device, and driver evidence before deciding on recovery or replacement.

Quick answer

Xid 79 reports that the driver cannot access a GPU over its PCIe connection. PCIe link failure is a common explanation, but GPU hardware and driver issues can also cause it. Preserve the evidence, drain affected work, then follow the platform restart procedure. The code alone does not prove that the card needs replacement or that a warm reboot cannot work.

Symptom
Xid 79: GPU has fallen off the bus
Root cause
A PCIe link failure can make the endpoint inaccessible. GPU hardware failure can produce the same event. Driver issues can also cause Xid 79; the event is not proof of an exclusively electrical fault.
Recommended fix
Preserve kernel logs, the event timestamp, GPU identity, and an NVIDIA bug report before disruptive recovery where possible.
How Denpex helps
Denpex matches Xid 79: GPU has fallen off the bus across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
Hardware#xid#xid-79#pcie#nvidia-driver#hardware#rma

What this failure is

NVIDIA emits Xid 79 when an attempted PCIe access finds the GPU inaccessible. It identifies a loss of device accessibility, not a complete diagnosis of the underlying component.

Live diagnosis, no signup

Is this what broke your run? Paste your log.

You're reading about Xid 79: GPU has fallen off the bus. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.

training_logs.txt
No log to hand? Try one:

3 free diagnoses/day

Why it happens (the mechanism)

NVIDIA identifies PCIe link failures, failing GPU hardware, and other driver issues as possible explanations. Correlate the event with kernel PCIe logs, platform events, driver version, and the affected device identity. Power and thermal records can help investigate a platform fault but do not establish one merely because Xid 79 occurred.

What you'll observe

  • A job loses access to a GPU
  • The driver reports that the GPU is no longer accessible over PCIe
  • The fault may recur after recovery and require vendor investigation

Common symptoms and what they mean

SymptomWhy it happens
NVRM: Xid with code 79 and the message GPU has fallen off the busA PCIe link failure can make the endpoint inaccessible
GPU enumeration or management commands may fail for the affected deviceGPU hardware failure can produce the same event
Kernel PCIe events may provide additional contextDriver issues can also cause Xid 79; the event is not proof of an exclusively electrical fault

Which systems are affected

  • NVIDIA GPUs and supported Linux drivers
  • Bare-metal GPU hosts and virtualized environments where host evidence is available

How to confirm this is the problem

Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.

  • ✓Locate Xid 79 in the host kernel log
  • ✓Record the affected PCI bus address and GPU identity
  • ✓Compare PCIe and system event logs around the incident
  • ✓Record recovery results and whether the problem recurs

Root cause

  • A PCIe link failure can make the endpoint inaccessible
  • GPU hardware failure can produce the same event
  • Driver issues can also cause Xid 79; the event is not proof of an exclusively electrical fault

The fix and how to prevent it

Searchable error signature

search key
Xid 79: GPU has fallen off the bus

Use this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.

The fix and the prevention pattern

The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.

Sign up free. Unlock the full analysis

No credit card. Daily allowance follows verified trust tier. Instant access.

Why the recommended fix works

Evidence collection preserves the distinction between a device-access failure and its cause. A coordinated restart addresses the immediate host condition while subsequent diagnostics and recurrence evidence guide the vendor investigation.

Code examples

snippet
# Read-only evidence collection. Run with the permissions required by your host.
dmesg -T | grep -iE 'NVRM|Xid|pcieport|AER'
nvidia-smi -L
lspci -nn
# Collect the vendor report before recovery where possible.
sudo nvidia-bug-report.sh
# Review the report for sensitive host details before sharing it.

Adapt the snippet to your framework. The same pattern holds for PyTorch Lightning, Hugging Face Trainer, DeepSpeed, Megatron-LM, and vLLM training wrappers. Where the wrapper exposes a config flag (for examplelr_scheduler_type in Trainer), prefer the flag over the imperative API to keep the schedule declarative and reproducible.

Best practices by model family

Model / StackRecommendationNotes
Bare-metal hostPreserve evidence and coordinate restartDrain work before disruptive recovery.
Cloud or shared hostEscalate to the host ownerGuest logs may not expose the PCIe or platform cause.

With the fix vs without the fix

DimensionWith the fixWithout the fix
Event versus causeXid 79 plus kernel, platform, and driver evidenceAssuming the Xid proves a failed card
RecoveryPlatform-supported restart and validationAn unconditional power-cycle or replacement instruction

Diagnostic note

“This guidance follows NVIDIA's Xid catalog. It is not a claim of a captured Denpex customer incident. The example is a representative error signature, not a complete incident log.”

Visual fingerprint

Evidence-led recovery
Xid 79 -> preserve logs -> drain work -> supported host recovery -> verify -> investigate recurrence
The Xid identifies lost device access; investigation determines the cause.

Xid 79 in context

Xid 79 is one of a small set of codes the NVIDIA driver uses to report GPU faults, and the number is most of the diagnosis: it tells you whether you are looking at your own code, the driver, or a board that needs replacing.

Compare every Xid code side by side

Diagnose this failure in VS Code

Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.

Install the free VS Code extension

Frequently asked questions

Questions engineers and on-call staff commonly ask about this failure.

Does Xid 79 prove that the GPU needs replacement?
No. NVIDIA lists PCIe link, GPU hardware, and driver issues. Replacement should follow the vendor investigation, not the code alone.
Is a full power cycle always required?
No universal requirement follows from the code. NVIDIA lists a bare-metal host restart; follow the specific platform recovery procedure, which may require a power cycle.
What should I collect before restarting?
Preserve the kernel event and surrounding PCIe messages, device identity, platform logs where available, and an NVIDIA bug report. Review sensitive details before sharing.

Don't just read the fix, diagnose your run

The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.