Skip to content

NVLink Link Down: Peer Evidence and Recovery Checks

Distinguish an inactive topology link, NVLink errors and a failed peer. Preserve link identity and workload evidence before approved recovery.

Quick answer

An inactive NVLink or recovery counter alone does not establish a noisy or defective link. Compare the expected topology, link identity, counters over time and peer Xid events. A peer failure can cause an otherwise healthy GPU to report a link problem.

Root cause
The link evidence must be correlated with expected topology and both endpoints. Xid 74 can be a consequence of the remote device failing rather than a defective local GPU.
Recommended fix
Capture expected topology, link identity and full local/peer Xid lines before changing device state.
How Denpex helps
Denpex investigates NVLink Link Down: Peer Evidence and Recovery Checks using the evidence you provide or your connected workload collects. Earlier rank, host or application evidence is needed to distinguish an initiating failure from a downstream report.
Hardware#nvlink#recovery#link#state#down#replay

What this failure is

NVLink state and counters describe observed link behavior. Interpret them against the expected platform topology and the driver event at the time of failure.

Live diagnosis, no signup

Is this what broke your run? Paste your log.

You're reading about NVLink Link Down: Peer Evidence and Recovery Checks. Paste your own traceback and relevant evidence for an investigation of your workload, with a next action or a specific missing fact. A reference entry does not establish your cause. No account or card for the free diagnosis. Review data handling before submitting sensitive logs.

Before uploading, review cloud data handling and local options.

training_logs.txt
No log to hand? Try one:

3 free diagnoses/day

Why it happens (the mechanism)

A connection or remote device can be affected, while some links may be inactive by design. The complete event and peer topology distinguish these cases.

What you'll observe

  • Collective or transfer behavior changes while NVLink state or error counters differ from the expected topology.

Common symptoms and what they mean

SymptomWhy it happens
nvidia-smi nvlink reports an inactive link or changing error countersThe link evidence must be correlated with expected topology and both endpoints. Xid 74 can be a consequence of the remote device failing rather than a defective local GPU.

Which systems are affected

  • NVIDIA platforms with NVLink connections or topology checks

How to confirm this is the problem

Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.

  • ✓Compare nvidia-smi topology and link status with the expected machine configuration.
  • ✓Correlate local and remote events, accounting for clock skew.
  • ✓Check affected collective completion, output correctness and recurrence after approved recovery.

Root cause

  • The link evidence must be correlated with expected topology and both endpoints. Xid 74 can be a consequence of the remote device failing rather than a defective local GPU.

The fix and how to prevent it

The fix and the prevention pattern

The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.

Sign up free. Unlock the full analysis

No credit card. Daily allowance follows verified trust tier. Instant access.

Why the recommended fix works

Identifying the initiating endpoint prevents a downstream reporter being mistaken for the defective component.

Code examples

snippet
nvidia-smi topo -m
nvidia-smi nvlink -s

Adapt the snippet to your framework. The same pattern holds for PyTorch Lightning, Hugging Face Trainer, DeepSpeed, Megatron-LM, and vLLM training wrappers. Where the wrapper exposes a config flag (for examplelr_scheduler_type in Trainer), prefer the flag over the imperative API to keep the schedule declarative and reproducible.

Best practices by model family

Model / StackRecommendationNotes
Small CNN / MLPRecommendedStabilises early-gradient noise even for tiny models.
Transformer (ViT/BERT)RequiredAttention stacks amplify gradient instability without active mitigation.
LLM (Llama / Qwen / GPT)RequiredAt scale, every failure compounds across distributed collectives.
Diffusion / Stable DiffusionRecommendedU-Net + cross-attention paths benefit from the same hardening.

With the fix vs without the fix

DimensionWith the fixWithout the fix
Symptom windowStable from the first stepVisible within tens to hundreds of steps
Final metricsReproducible optimaPlateau or divergence below the baseline
Operational riskBounded by the prevention checklistCompounds across folds / reruns
Prod recommendationShipBlock until the fix is in place

Diagnostic note

“This is an evidence method, not a claim of tested physical NVLink recovery on the single-GPU replay machine.”

Diagnose this failure in VS Code

Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.

Install the free VS Code extension

Frequently asked questions

Questions engineers and on-call staff commonly ask about this failure.

Is every inactive link a fault?
No. Compare it with the expected topology and the workload path.
Does Xid 74 prove the reporting GPU is bad?
No. It can follow a remote device or link failure. Inspect both endpoints before assigning ownership.

Don't just read the fix, diagnose your run

The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.