Skip to content

Xid 92: High single-bit ECC error rate

Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss.

Quick answer

Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss.

Symptom
Xid 92
Root cause
Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss.
Recommended fix
Record volatile and aggregate ECC counters, timestamps, GPU UUID, and row-remapping status. Compare counter deltas over time and schedule vendor-supported diagnostics after draining work. Investigate NaNs independently rather than attributing them to corrected bits.
How Denpex helps
Denpex investigates Xid 92: High single-bit ECC error rate using the evidence you provide or your connected workload collects. Earlier rank, host or application evidence is needed to distinguish an initiating failure from a downstream report.
Hardware#xid#nvidia#corrected-ecc#gpu

What this failure is

Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss.

Live diagnosis, no signup

Is this what broke your run? Paste your log.

You're reading about Xid 92: High single-bit ECC error rate. Paste your own traceback and relevant evidence for an investigation of your workload, with a next action or a specific missing fact. A reference entry does not establish your cause. No account or card for the free diagnosis. Review data handling before submitting sensitive logs.

Before uploading, review cloud data handling and local options.

training_logs.txt
No log to hand? Try one:

3 free diagnoses/day

Why it happens (the mechanism)

Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss. Use incident telemetry and the supported driver/firmware baseline to investigate the underlying mechanism.

What you'll observe

  • Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss.
  • The event requires investigation; it does not establish every proposed downstream effect.

Common symptoms and what they mean

SymptomWhy it happens
Xid 92: High single-bit ECC error rateXid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss.
Record any accompanying application failure separately from this driver event.Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss.

Which systems are affected

  • NVIDIA data center GPUs and the Linux driver
  • production-shaped multi-accelerator workloads
  • containerized and bare-metal deployments of the same stack

How to confirm this is the problem

Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.

  • ✓Preserve the literal Xid code and incident timestamp from kernel logs.
  • ✓Record nvidia-smi -q output, including GPU UUID, ECC and row-remapping details.
  • ✓Compare fresh counter deltas and diagnostics with the fleet baseline.

Root cause

  • Xid 92 reports a high rate of corrected single-bit ECC errors. Corrected errors preserve the original data; this event does not prove uncorrectable memory errors, corrupted tensors, or the cause of NaN loss.

The fix and how to prevent it

Searchable error signature

search key
Xid 92

Use this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.

The fix and the prevention pattern

The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.

Sign up free. Unlock the full analysis

No credit card. Daily allowance follows verified trust tier. Instant access.

Why the recommended fix works

Trend monitoring and diagnostics distinguish a transient correction from persistent degradation without assuming data corruption.

Code examples

snippet
# Read-only evidence collection
dmesg -T | grep -iE "NVRM|Xid|ECC"
nvidia-smi -q -d ECC,ROW_REMAPPER
# Preserve a GPU bug report before reset or reboot
sudo nvidia-bug-report.sh

Adapt the snippet to your framework. The same pattern holds for PyTorch Lightning, Hugging Face Trainer, DeepSpeed, Megatron-LM, and vLLM training wrappers. Where the wrapper exposes a config flag (for examplelr_scheduler_type in Trainer), prefer the flag over the imperative API to keep the schedule declarative and reproducible.

Best practices by model family

Model / StackRecommendationNotes
EvidenceKeep raw logsRecord GPU identity, timestamps and counter deltas.
RecoveryFollow the platform procedureStop affected work before disruptive diagnostics or reset.

With the fix vs without the fix

DimensionWith the fixWithout the fix
Event versus causeLiteral Xid plus supporting telemetryUnsupported hardware or workload attribution

Diagnostic note

“This guidance follows NVIDIA Xid documentation. A documented event meaning is not proof of permanent hardware failure or a diagnosis of application-level NaNs.”

Visual fingerprint

Evidence-led response
capture event -> identify GPU -> inspect counters and context -> supported recovery or diagnostics -> validate
Escalate unresolved or recurring faults with collected evidence.

Xid 92 in context

Xid 92 is one of a small set of codes the NVIDIA driver uses to report GPU faults, and the number is most of the diagnosis: it tells you whether you are looking at your own code, the driver, or a board that needs replacing.

Compare every Xid code side by side

Diagnose this failure in VS Code

Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.

Install the free VS Code extension

Frequently asked questions

Questions engineers and on-call staff commonly ask about this failure.

Does this prove permanent hardware failure?
No. The code identifies an event; vendor investigation and diagnostics determine the underlying fault.
What should I do first?
Record volatile and aggregate ECC counters, timestamps, GPU UUID, and row-remapping status. Compare counter deltas over time and schedule vendor-supported diagnostics after draining work. Investigate NaNs independently rather than attributing them to corrected bits.
When should I replace the GPU?
Escalate recurrence, rising error rates, exhausted remap capacity, or failed diagnostics to NVIDIA/OEM support. Replacement follows vendor investigation; no universal two-event or 30-day RMA rule is established here.

Don't just read the fix, diagnose your run

The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.