Xid 43: Watchdog Timeout due to PSU Transient Spikes
During heavy compute operations, modern GPUs (like the A100 or H100) experience massive microsecond-level power spikes. If the Power Supply Unit (PSU) or the 12V PCIe power cables cannot handle the transient load, the GPU voltage drops. This physical brownout stalls the GPU clocks, tripping the internal driver watchdog timer, which reports an Xid 43 (stopped processing) and resets the channel.
During heavy compute operations, modern GPUs (like the A100 or H100) experience massive microsecond-level power spikes.
- Symptom
The node may completely freeze, or the affected GPU may fall off the bus entirely (leading to Xid 79).- Root cause
- During heavy compute operations, modern GPUs (like the A100 or H100) experience massive microsecond-level power spikes. If the Power Supply Unit (PSU) or the 12V PCIe power cables cannot handle the transient load, the GPU voltage drops. This physical brownout stalls the GPU clocks, tripping the internal driver watchdog timer, which reports an Xid 43 (stopped processing) and resets the channel.
- Recommended fix
- Limit Maximum Power Draw sudo nvidia-smi -pl 250 Lowers the GPU power limit (e.g., to 250W), reducing the magnitude of power spikes and keeping the system stable until the PSU or cabling is upgraded.
- How Denpex helps
- Denpex matches Xid 43: Watchdog Timeout due to PSU Transient Spikes across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
What this failure is
Xid 43: Watchdog Timeout due to PSU Transient Spikes is a Hardware failure seen during ML training runs. During heavy compute operations, modern GPUs (like the A100 or H100) experience massive microsecond-level power spikes. If the Power Supply Unit (PSU) or the 12V PCIe power cables cannot handle the transient load, the GPU voltage drops. This physical brownout stalls the GPU clocks, tripping the internal driver watchdog timer, which reports an Xid 43 (stopped processing) and resets the channel. Common tags: Xid Error.
Is this what broke your run? Paste your log.
You're reading about Xid 43: Watchdog Timeout due to PSU Transient Spikes. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.
Want 14 days on the Scale plan?
Request an evaluation code. A verified workplace organization activates up to 50 diagnoses a day, alerts, history, and follow-up questions. No credit card or automatic subscription.
Why it happens (the mechanism)
Because the software reports a channel reset or watchdog timeout, engineers assume their code got stuck in an infinite loop. They spend weeks profiling CUDA kernels when the issue is actually an inadequate power supply.
What you'll observe
- NVRM: Xid (PCI:0000:81:00): 43, Ch 00000008, engmask 00000101
- GPU stopped processing
- Reset Channel Verification Error
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| Under high load (e.g., LLM training or heavy batch inference), a specific GPU stops responding. | During heavy compute operations, modern GPUs (like the A100 or H100) experience massive microsecond-level power spikes. If the Power Supply Unit (PSU) or the 12V PCIe power cables cannot handle the transient load, the GPU voltage drops. This physical brownout stalls the GPU clocks, tripping the internal driver watchdog timer, which reports an Xid 43 (stopped processing) and resets the channel. |
| The node may completely freeze, or the affected GPU may fall off the bus entirely (leading to Xid 79). | During heavy compute operations, modern GPUs (like the A100 or H100) experience massive microsecond-level power spikes. If the Power Supply Unit (PSU) or the 12V PCIe power cables cannot handle the transient load, the GPU voltage drops. This physical brownout stalls the GPU clocks, tripping the internal driver watchdog timer, which reports an Xid 43 (stopped processing) and resets the channel. |
| nvidia-smi hangs or reports the GPU as lost. | During heavy compute operations, modern GPUs (like the A100 or H100) experience massive microsecond-level power spikes. If the Power Supply Unit (PSU) or the 12V PCIe power cables cannot handle the transient load, the GPU voltage drops. This physical brownout stalls the GPU clocks, tripping the internal driver watchdog timer, which reports an Xid 43 (stopped processing) and resets the channel. |
Which systems are affected
- Server Chassis
- PSU
- NVIDIA GPU
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Check power connections and ensure separate PCIe cables are used (no daisy-chaining/splitters).
- ✓Limit GPU clocks to artificially reduce power transients: sudo nvidia-smi -lgc 300,1200 and see if the crash stops.
- ✓Check baseboard management controller (BMC) or IPMI logs for power supply voltage drops.
Searchable error signature
The node may completely freeze, or the affected GPU may fall off the bus entirely (leading to Xid 79).
nvidia-smi hangs or reports the GPU as lost.Use this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Xid 43 in context
Xid 43 is one of a small set of codes the NVIDIA driver uses to report GPU faults, and the number is most of the diagnosis: it tells you whether you are looking at your own code, the driver, or a board that needs replacing.
Compare every Xid code side by sideDiagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionRoot cause
- During heavy compute operations, modern GPUs (like the A100 or H100) experience massive microsecond-level power spikes. If the Power Supply Unit (PSU) or the 12V PCIe power cables cannot handle the transient load, the GPU voltage drops. This physical brownout stalls the GPU clocks, tripping the internal driver watchdog timer, which reports an Xid 43 (stopped processing) and resets the channel.
The fix and how to prevent it
Evaluate Denpex on your own logs
Request a Scale evaluation code. A verified workplace organization activates 14 days with up to 50 diagnoses a day. Every account keeps its current diagnosis allowance and gets a verification path. No card or automatic subscription.
References
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.