Xid 92: High single-bit ECC error rate
Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss.
Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss.
- Symptom
Xid 92- Root cause
- Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss.
- Recommended fix
- Record volatile and aggregate ECC counters, timestamps, GPU UUID, and row-remapping status. Compare counter deltas over time and schedule vendor-supported diagnostics after draining work. Investigate NaNs independently rather than attributing them to corrected bits.
- How Denpex helps
- Denpex investigates Xid 92: High single-bit ECC error rate using the evidence you provide or your connected workload collects. Earlier rank, host or application evidence is needed to distinguish an initiating failure from a downstream report.
What this failure is
Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss.
Is this what broke your run? Paste your log.
You're reading about Xid 92: High single-bit ECC error rate. Paste your own traceback and relevant evidence for an investigation of your workload, with a next action or a specific missing fact. A reference entry does not establish your cause. No account or card for the free diagnosis. Review data handling before submitting sensitive logs.
Before uploading, review cloud data handling and local options.
Why it happens (the mechanism)
Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss. Use incident telemetry and the supported driver/firmware baseline to investigate the underlying mechanism.
What you'll observe
- Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss.
- The event requires investigation; it does not establish every proposed downstream effect.
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| Xid 92: High single-bit ECC error rate | Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss. |
| Record any accompanying application failure separately from this driver event. | Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss. |
Which systems are affected
- NVIDIA data center GPUs and the Linux driver
- production-shaped multi-accelerator workloads
- containerized and bare-metal deployments of the same stack
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Preserve the literal Xid code and incident timestamp from kernel logs.
- ✓Record nvidia-smi -q output, including GPU UUID, ECC and row-remapping details.
- ✓Compare fresh counter deltas and diagnostics with the fleet baseline.
Root cause
- Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss.
The fix and how to prevent it
Searchable error signature
Xid 92Use this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Why the recommended fix works
Trend monitoring and diagnostics distinguish a transient correction from persistent degradation without assuming data corruption.
Code examples
# Read-only evidence collection
dmesg -T | grep -iE "NVRM|Xid|ECC"
nvidia-smi -q -d ECC,ROW_REMAPPER
# Preserve a GPU bug report before reset or reboot
sudo nvidia-bug-report.shAdapt the snippet to your framework. The same pattern holds for PyTorch Lightning, Hugging Face Trainer, DeepSpeed, Megatron-LM, and vLLM training wrappers. Where the wrapper exposes a config flag (for examplelr_scheduler_type in Trainer), prefer the flag over the imperative API to keep the schedule declarative and reproducible.
Best practices by model family
| Model / Stack | Recommendation | Notes |
|---|---|---|
| Evidence | Keep raw logs | Record GPU identity, timestamps and counter deltas. |
| Recovery | Follow the platform procedure | Stop affected work before disruptive diagnostics or reset. |
With the fix vs without the fix
| Dimension | With the fix | Without the fix |
|---|---|---|
| Event versus cause | Literal Xid plus supporting telemetry | Unsupported hardware or workload attribution |
Diagnostic note
“This guidance follows NVIDIA Xid documentation. A documented event meaning is not proof of permanent hardware failure or a diagnosis of application-level NaNs.”
Visual fingerprint
capture event -> identify GPU -> inspect counters and context -> supported recovery or diagnostics -> validate
Xid 92 in context
Xid 92 is one of a small set of codes the NVIDIA driver uses to report GPU faults, and the number is most of the diagnosis: it tells you whether you are looking at your own code, the driver, or a board that needs replacing.
Compare every Xid code side by sideDiagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionRelated failures to investigate next
Frequently asked questions
Questions engineers and on-call staff commonly ask about this failure.
Does this prove permanent hardware failure?
What should I do first?
When should I replace the GPU?
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.