Xid 62: Internal microcontroller halt
Xid 62 reports an internal microcontroller halt. It does not by itself distinguish a driver or firmware defect from faulty hardware, and does not establish a CUDA memory leak.
Xid 62 reports an internal microcontroller halt. It does not by itself distinguish a driver or firmware defect from faulty hardware, and does not establish a CUDA memory leak.
- Symptom
NVRM: Xid (PCI:0000:41:00.0): 62, pid=14092, 0000 00000000 00000000- Root cause
- Xid 62 reports an internal microcontroller halt. It does not by itself distinguish a driver or firmware defect from faulty hardware, and does not establish a CUDA memory leak.
- Recommended fix
- Preserve nvidia-bug-report.sh output and kernel logs, drain work and stop GPU clients, then attempt a supported nvidia-smi --gpu-reset -i <id>. If reset fails, follow the platform reboot or power-cycle procedure. Validate the GPU before returning it to service.
- How Denpex helps
- Denpex investigates Xid 62: Internal microcontroller halt using the evidence you provide or your connected workload collects. Earlier rank, host or application evidence is needed to distinguish an initiating failure from a downstream report.
What this failure is
Xid 62 reports an internal microcontroller halt. It does not by itself distinguish a driver or firmware defect from faulty hardware, and does not establish a CUDA memory leak.
Is this what broke your run? Paste your log.
You're reading about Xid 62: Internal microcontroller halt. Paste your own traceback and relevant evidence for an investigation of your workload, with a next action or a specific missing fact. A reference entry does not establish your cause. No account or card for the free diagnosis. Review data handling before submitting sensitive logs.
Before uploading, review cloud data handling and local options.
Why it happens (the mechanism)
Xid 62 reports an internal microcontroller halt. It does not by itself distinguish a driver or firmware defect from faulty hardware, and does not establish a CUDA memory leak. Use incident telemetry and the supported driver/firmware baseline to investigate the underlying mechanism.
What you'll observe
- Xid 62 reports an internal microcontroller halt. It does not by itself distinguish a driver or firmware defect from faulty hardware, and does not establish a CUDA memory leak.
- The event requires investigation; it does not establish every proposed downstream effect.
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| Xid 62: Internal microcontroller halt | Xid 62 reports an internal microcontroller halt. It does not by itself distinguish a driver or firmware defect from faulty hardware, and does not establish a CUDA memory leak. |
| Record any accompanying application failure separately from this driver event. | Xid 62 reports an internal microcontroller halt. It does not by itself distinguish a driver or firmware defect from faulty hardware, and does not establish a CUDA memory leak. |
Which systems are affected
- NVIDIA data center GPUs and the Linux driver
- production-shaped multi-accelerator workloads
- containerized and bare-metal deployments of the same stack
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Preserve the literal Xid code and incident timestamp from kernel logs.
- ✓Record nvidia-smi -q output, including GPU UUID, ECC and row-remapping details.
- ✓Compare fresh counter deltas and diagnostics with the fleet baseline.
Root cause
- Xid 62 reports an internal microcontroller halt. It does not by itself distinguish a driver or firmware defect from faulty hardware, and does not establish a CUDA memory leak.
The fix and how to prevent it
Searchable error signature
NVRM: Xid (PCI:0000:41:00.0): 62, pid=14092, 0000 00000000 00000000Use this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Why the recommended fix works
A supported reset may restore the halted controller. Successful recovery does not prove that the underlying cause is fixed.
Code examples
# Read-only evidence collection
dmesg -T | grep -iE "NVRM|Xid|ECC"
nvidia-smi -q -d ECC,ROW_REMAPPER
# Preserve a GPU bug report before reset or reboot
sudo nvidia-bug-report.shAdapt the snippet to your framework. The same pattern holds for PyTorch Lightning, Hugging Face Trainer, DeepSpeed, Megatron-LM, and vLLM training wrappers. Where the wrapper exposes a config flag (for examplelr_scheduler_type in Trainer), prefer the flag over the imperative API to keep the schedule declarative and reproducible.
Best practices by model family
| Model / Stack | Recommendation | Notes |
|---|---|---|
| Evidence | Keep raw logs | Record GPU identity, timestamps and counter deltas. |
| Recovery | Follow the platform procedure | Stop affected work before disruptive diagnostics or reset. |
With the fix vs without the fix
| Dimension | With the fix | Without the fix |
|---|---|---|
| Event versus cause | Literal Xid plus supporting telemetry | Unsupported hardware or workload attribution |
Diagnostic note
“This guidance follows NVIDIA Xid documentation. A documented event meaning is not proof of permanent hardware failure or a diagnosis of application-level NaNs.”
Visual fingerprint
capture event -> identify GPU -> inspect counters and context -> supported recovery or diagnostics -> validate
Xid 62 in context
Xid 62 is one of a small set of codes the NVIDIA driver uses to report GPU faults, and the number is most of the diagnosis: it tells you whether you are looking at your own code, the driver, or a board that needs replacing.
Compare every Xid code side by sideDiagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionRelated failures to investigate next
Frequently asked questions
Questions engineers and on-call staff commonly ask about this failure.
Does this prove permanent hardware failure?
What should I do first?
When should I replace the GPU?
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.