Skip to content

A worker loses its GPU mid-run and the job hangs until a collective times out

When a GPU disappears from its host, a bus fault, a driver reset, a thermal or power event, the rank owning it stops answering. Its peers are inside a collective waiting for data that will never arrive, so the job does not crash at the moment of failure. It stalls, and the first error anyone sees is a timeout on a healthy rank several minutes later.

Quick answer

The rank in the error message is a witness, not the culprit. Take the time of the last progress line and search every node's kernel log at that moment for an accelerator fault.

Symptom
RuntimeError: Worker failed with error 'level_zero backend failed with error: 20 (UR_RESULT_ERROR_DEVICE_LOST)'
Root cause
A collective is a rendezvous. Every participant blocks until all of them arrive, and there is no mechanism by which a waiting rank learns that a peer has ceased to exist, it can only observe that the data has not come. The absence of a peer and a peer that is merely slow are indistinguishable until a deadline expires.
Recommended fix
Take the timestamp of the last normal progress line, not the timestamp of the error, and look in the host kernel log of every node at that moment. The device-loss event is recorded there when it happens.
How Denpex helps
Denpex matches A worker loses its GPU mid-run and the job hangs until a collective times out across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
Reliability#device-lost#xid-79#hang#collective-timeout#pipeline-parallel#level-zero

What this failure is

A stall in which one rank's accelerator becomes unreachable, leaving its peers blocked inside a collective until a timeout fires, so the failure is reported by a healthy rank long after and elsewhere from the fault.

Live diagnosis, no signup

Is this what broke your run? Paste your log.

You're reading about A worker loses its GPU mid-run and the job hangs until a collective times out. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.

training_logs.txt
No log to hand? Try one:

3 free diagnoses/day

Why it happens (the mechanism)

Collectives have no liveness channel. A rank waiting for data cannot distinguish a peer that is slow from a peer that no longer exists, so it waits for its deadline and then reports what it experienced, which is a timeout. The rank that actually failed lost the device it would have needed to report with, so it says nothing. Every visible piece of evidence therefore comes from the wrong place and the wrong time.

What you'll observe

  • The reported error names a rank that is working correctly, not the one that failed
  • Minutes pass between the real fault and any log line, so the timestamps mislead
  • The job looks alive to the scheduler throughout the stall and keeps its allocation
  • Restarting sometimes works, which suggests a transient software fault rather than hardware

Common symptoms and what they mean

SymptomWhy it happens
A fatal timeout arriving several minutes after the last normal progress lineA collective is a rendezvous. Every participant blocks until all of them arrive, and there is no mechanism by which a waiting rank learns that a peer has ceased to exist, it can only observe that the data has not come. The absence of a peer and a peer that is merely slow are indistinguishable until a deadline expires.
RuntimeError reporting a worker failed with a device-lost error from the runtime backendSo the failure inverts the usual relationship between cause and report. The rank whose device vanished cannot report anything, because the thing it would report with is gone. The report comes from a healthy rank, describing a symptom of someone else's fault, at a time chosen by the timeout rather than by the event.
level_zero backend failed with error 20, UR_RESULT_ERROR_DEVICE_LOST on Intel acceleratorsThe device loss itself is a host-level event. It is recorded in the kernel log at the instant it happens, which is why that log and the job's own output disagree about when the failure occurred, often by the whole length of the timeout.
Xid 79 or a GPU has fallen off the bus message in the host kernel log at the moment progress stoppedA collective is a rendezvous. Every participant blocks until all of them arrive, and there is no mechanism by which a waiting rank learns that a peer has ceased to exist, it can only observe that the data has not come. The absence of a peer and a peer that is merely slow are indistinguishable until a deadline expires.
Peer ranks reporting a collective timeout while the failed rank reports nothing at allSo the failure inverts the usual relationship between cause and report. The rank whose device vanished cannot report anything, because the thing it would report with is gone. The report comes from a healthy rank, describing a symptom of someone else's fault, at a time chosen by the timeout rather than by the event.

Which systems are affected

  • Pipeline-parallel and tensor-parallel jobs, where one lost rank stalls every rank downstream of it
  • Inference engines that fan a step out to workers and wait on the result
  • NVIDIA GPUs raising Xid 79, and Intel accelerators raising a level-zero device-lost result
  • Any collective-based workload, since the waiting side is what reports the failure

How to confirm this is the problem

Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.

  • ✓Compare the timestamp of the last progress line against the timestamp of the error. A gap matching the configured timeout means the job was stalled, not working, for that interval.
  • ✓Search each node's kernel log for an accelerator fault at the time progress stopped; that entry names the failing device directly.
  • ✓Determine which rank fell silent first from the per-rank output. That rank, not the one that raised the timeout, is where the fault is.

Root cause

  • A collective is a rendezvous. Every participant blocks until all of them arrive, and there is no mechanism by which a waiting rank learns that a peer has ceased to exist, it can only observe that the data has not come. The absence of a peer and a peer that is merely slow are indistinguishable until a deadline expires.
  • So the failure inverts the usual relationship between cause and report. The rank whose device vanished cannot report anything, because the thing it would report with is gone. The report comes from a healthy rank, describing a symptom of someone else's fault, at a time chosen by the timeout rather than by the event.
  • The device loss itself is a host-level event. It is recorded in the kernel log at the instant it happens, which is why that log and the job's own output disagree about when the failure occurred, often by the whole length of the timeout.

The fix and how to prevent it

Searchable error signature

search key
RuntimeError: Worker failed with error 'level_zero backend failed with error: 20 (UR_RESULT_ERROR_DEVICE_LOST)'
NVRM: Xid (PCI:0000:41:00): 79, pid=..., GPU has fallen off the bus.
[rank2] Watchdog caught collective operation timeout: WorkNCCL(SeqNum=8412, OpType=ALLREDUCE) ran for 1800000 milliseconds before timing out

Use this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.

The fix and the prevention pattern

The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.

Sign up free. Unlock the full analysis

No credit card. Daily allowance follows verified trust tier. Instant access.

Why the recommended fix works

Anchoring the investigation to the last progress line rather than the error moves it back to when the fault happened, and the host kernel log at that moment records the device event directly. That converts an unattributable job-level timeout into a named piece of hardware, which is the only form of the problem that can be acted on.

Best practices by model family

Model / StackRecommendationNotes
Timeout reported on one rankFind the rank that fell silent firstThe reporting rank is a witness to someone else's failure.
Investigating after the factSearch host kernel logs at the last-progress timestampThe device event is recorded when it happens, not when the job notices.
Suspected node identifiedDrain and test rather than restart onto itA device lost once is likely to be lost again.
Diagnosing activelyShorten the collective timeout temporarilyBrings the report closer to the event; restore it afterwards.

With the fix vs without the fix

DimensionWith the fixWithout the fix
Which rank the error namesA healthy rank that was waitingAssumed to be the rank that failed
When the error appearsOne timeout interval after the faultRead as the time of the failure
What the job does meanwhileHolds its allocation and looks aliveExpected to crash promptly

Diagnostic note

“The most expensive mistake here is restarting onto the same pool. A restart usually places ranks differently, so the job survives, the incident is closed, and a device that has already failed once stays in service until it takes down something bigger. Treat a single confirmed device loss as grounds to drain the node, not as a transient to retry through.”

Visual fingerprint

The fault and the report are on different ranks and different clocks
  t+0     rank 3 device lost        kernel log records it   (no job output)
  t+0     ranks 0,1,2 enter collective and block
  ...     job holds its allocation, scheduler sees it as running
  t+30m   rank 2 watchdog fires   -> collective timeout reported HERE

  visible evidence: rank 2, at t+30m.   actual fault: rank 3, at t+0.
The rank whose device disappeared produces no output because the device it would report with is gone. A peer blocked in the collective reports a timeout when its deadline expires, so both the rank named and the time reported differ from the fault.

Xid 79 in context

Xid 79 is one of a small set of codes the NVIDIA driver uses to report GPU faults, and the number is most of the diagnosis: it tells you whether you are looking at your own code, the driver, or a board that needs replacing.

Compare every Xid code side by side

Diagnose this failure in VS Code

Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.

Install the free VS Code extension

Frequently asked questions

Questions engineers and on-call staff commonly ask about this failure.

The error names rank 2, so why look at rank 3?
Rank 2 was waiting in a collective and reported a timeout. The rank that failed lost the device it would have used to report, so it is silent. Find the first rank to go quiet.
Why did nothing fail for half an hour?
Collectives block rather than fail. Nothing detects a missing peer until the timeout expires, so the gap between fault and report is the timeout interval.
A restart fixed it. Was it transient?
Usually not. A restart tends to place ranks on different hardware, so the job avoids the faulty device rather than the device being repaired.
Is this the same as a straggler?
It produces the same waiting, which is why a timeout cannot distinguish them. A straggler eventually arrives; a lost device never does.

Don't just read the fix, diagnose your run

The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.