A worker loses its GPU mid-run and the job hangs until a collective times out
When a GPU disappears from its host, a bus fault, a driver reset, a thermal or power event, the rank owning it stops answering. Its peers are inside a collective waiting for data that will never arrive, so the job does not crash at the moment of failure. It stalls, and the first error anyone sees is a timeout on a healthy rank several minutes later.
The rank in the error message is a witness, not the culprit. Take the time of the last progress line and search every node's kernel log at that moment for an accelerator fault.
- Symptom
RuntimeError: Worker failed with error 'level_zero backend failed with error: 20 (UR_RESULT_ERROR_DEVICE_LOST)'- Root cause
- A collective is a rendezvous. Every participant blocks until all of them arrive, and there is no mechanism by which a waiting rank learns that a peer has ceased to exist, it can only observe that the data has not come. The absence of a peer and a peer that is merely slow are indistinguishable until a deadline expires.
- Recommended fix
- Take the timestamp of the last normal progress line, not the timestamp of the error, and look in the host kernel log of every node at that moment. The device-loss event is recorded there when it happens.
- How Denpex helps
- Denpex matches A worker loses its GPU mid-run and the job hangs until a collective times out across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
What this failure is
A stall in which one rank's accelerator becomes unreachable, leaving its peers blocked inside a collective until a timeout fires, so the failure is reported by a healthy rank long after and elsewhere from the fault.
Is this what broke your run? Paste your log.
You're reading about A worker loses its GPU mid-run and the job hangs until a collective times out. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.
Why it happens (the mechanism)
Collectives have no liveness channel. A rank waiting for data cannot distinguish a peer that is slow from a peer that no longer exists, so it waits for its deadline and then reports what it experienced, which is a timeout. The rank that actually failed lost the device it would have needed to report with, so it says nothing. Every visible piece of evidence therefore comes from the wrong place and the wrong time.
What you'll observe
- The reported error names a rank that is working correctly, not the one that failed
- Minutes pass between the real fault and any log line, so the timestamps mislead
- The job looks alive to the scheduler throughout the stall and keeps its allocation
- Restarting sometimes works, which suggests a transient software fault rather than hardware
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| A fatal timeout arriving several minutes after the last normal progress line | A collective is a rendezvous. Every participant blocks until all of them arrive, and there is no mechanism by which a waiting rank learns that a peer has ceased to exist, it can only observe that the data has not come. The absence of a peer and a peer that is merely slow are indistinguishable until a deadline expires. |
| RuntimeError reporting a worker failed with a device-lost error from the runtime backend | So the failure inverts the usual relationship between cause and report. The rank whose device vanished cannot report anything, because the thing it would report with is gone. The report comes from a healthy rank, describing a symptom of someone else's fault, at a time chosen by the timeout rather than by the event. |
| level_zero backend failed with error 20, UR_RESULT_ERROR_DEVICE_LOST on Intel accelerators | The device loss itself is a host-level event. It is recorded in the kernel log at the instant it happens, which is why that log and the job's own output disagree about when the failure occurred, often by the whole length of the timeout. |
| Xid 79 or a GPU has fallen off the bus message in the host kernel log at the moment progress stopped | A collective is a rendezvous. Every participant blocks until all of them arrive, and there is no mechanism by which a waiting rank learns that a peer has ceased to exist, it can only observe that the data has not come. The absence of a peer and a peer that is merely slow are indistinguishable until a deadline expires. |
| Peer ranks reporting a collective timeout while the failed rank reports nothing at all | So the failure inverts the usual relationship between cause and report. The rank whose device vanished cannot report anything, because the thing it would report with is gone. The report comes from a healthy rank, describing a symptom of someone else's fault, at a time chosen by the timeout rather than by the event. |
Which systems are affected
- Pipeline-parallel and tensor-parallel jobs, where one lost rank stalls every rank downstream of it
- Inference engines that fan a step out to workers and wait on the result
- NVIDIA GPUs raising Xid 79, and Intel accelerators raising a level-zero device-lost result
- Any collective-based workload, since the waiting side is what reports the failure
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Compare the timestamp of the last progress line against the timestamp of the error. A gap matching the configured timeout means the job was stalled, not working, for that interval.
- ✓Search each node's kernel log for an accelerator fault at the time progress stopped; that entry names the failing device directly.
- ✓Determine which rank fell silent first from the per-rank output. That rank, not the one that raised the timeout, is where the fault is.
Root cause
- A collective is a rendezvous. Every participant blocks until all of them arrive, and there is no mechanism by which a waiting rank learns that a peer has ceased to exist, it can only observe that the data has not come. The absence of a peer and a peer that is merely slow are indistinguishable until a deadline expires.
- So the failure inverts the usual relationship between cause and report. The rank whose device vanished cannot report anything, because the thing it would report with is gone. The report comes from a healthy rank, describing a symptom of someone else's fault, at a time chosen by the timeout rather than by the event.
- The device loss itself is a host-level event. It is recorded in the kernel log at the instant it happens, which is why that log and the job's own output disagree about when the failure occurred, often by the whole length of the timeout.
The fix and how to prevent it
Searchable error signature
RuntimeError: Worker failed with error 'level_zero backend failed with error: 20 (UR_RESULT_ERROR_DEVICE_LOST)'
NVRM: Xid (PCI:0000:41:00): 79, pid=..., GPU has fallen off the bus.
[rank2] Watchdog caught collective operation timeout: WorkNCCL(SeqNum=8412, OpType=ALLREDUCE) ran for 1800000 milliseconds before timing outUse this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Why the recommended fix works
Anchoring the investigation to the last progress line rather than the error moves it back to when the fault happened, and the host kernel log at that moment records the device event directly. That converts an unattributable job-level timeout into a named piece of hardware, which is the only form of the problem that can be acted on.
Best practices by model family
| Model / Stack | Recommendation | Notes |
|---|---|---|
| Timeout reported on one rank | Find the rank that fell silent first | The reporting rank is a witness to someone else's failure. |
| Investigating after the fact | Search host kernel logs at the last-progress timestamp | The device event is recorded when it happens, not when the job notices. |
| Suspected node identified | Drain and test rather than restart onto it | A device lost once is likely to be lost again. |
| Diagnosing actively | Shorten the collective timeout temporarily | Brings the report closer to the event; restore it afterwards. |
With the fix vs without the fix
| Dimension | With the fix | Without the fix |
|---|---|---|
| Which rank the error names | A healthy rank that was waiting | Assumed to be the rank that failed |
| When the error appears | One timeout interval after the fault | Read as the time of the failure |
| What the job does meanwhile | Holds its allocation and looks alive | Expected to crash promptly |
Diagnostic note
“The most expensive mistake here is restarting onto the same pool. A restart usually places ranks differently, so the job survives, the incident is closed, and a device that has already failed once stays in service until it takes down something bigger. Treat a single confirmed device loss as grounds to drain the node, not as a transient to retry through.”
Visual fingerprint
t+0 rank 3 device lost kernel log records it (no job output) t+0 ranks 0,1,2 enter collective and block ... job holds its allocation, scheduler sees it as running t+30m rank 2 watchdog fires -> collective timeout reported HERE visible evidence: rank 2, at t+30m. actual fault: rank 3, at t+0.
Xid 79 in context
Xid 79 is one of a small set of codes the NVIDIA driver uses to report GPU faults, and the number is most of the diagnosis: it tells you whether you are looking at your own code, the driver, or a board that needs replacing.
Compare every Xid code side by sideDiagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionRelated failures to investigate next
Frequently asked questions
Questions engineers and on-call staff commonly ask about this failure.
The error names rank 2, so why look at rank 3?
Why did nothing fail for half an hour?
A restart fixed it. Was it transient?
Is this the same as a straggler?
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.
Related Reliability errors
GB200 NVL72 driver hang with Xid 150 and 154 after long uptime is a firmware bug, not a dead GPU
Reliability · high
Sequence Length Imbalance Causing Distributed Training Stragglers
Reliability · high
Silent Data Corruption from GPU Hardware Faults Causing Loss Spikes and Model Divergence
Reliability · critical
NIXL Firmware Page Registration Fan-Out Triggers Host OOM Kills on HGX H200 and B200
Reliability · critical