NVLink Link Down: Peer Evidence and Recovery Checks
Distinguish an inactive topology link, NVLink errors and a failed peer. Preserve link identity and workload evidence before approved recovery.
An inactive NVLink or recovery counter alone does not establish a noisy or defective link. Compare the expected topology, link identity, counters over time and peer Xid events. A peer failure can cause an otherwise healthy GPU to report a link problem.
- Root cause
- The link evidence must be correlated with expected topology and both endpoints. Xid 74 can be a consequence of the remote device failing rather than a defective local GPU.
- Recommended fix
- Capture expected topology, link identity and full local/peer Xid lines before changing device state.
- How Denpex helps
- Denpex investigates NVLink Link Down: Peer Evidence and Recovery Checks using the evidence you provide or your connected workload collects. Earlier rank, host or application evidence is needed to distinguish an initiating failure from a downstream report.
What this failure is
NVLink state and counters describe observed link behavior. Interpret them against the expected platform topology and the driver event at the time of failure.
Is this what broke your run? Paste your log.
You're reading about NVLink Link Down: Peer Evidence and Recovery Checks. Paste your own traceback and relevant evidence for an investigation of your workload, with a next action or a specific missing fact. A reference entry does not establish your cause. No account or card for the free diagnosis. Review data handling before submitting sensitive logs.
Before uploading, review cloud data handling and local options.
Why it happens (the mechanism)
A connection or remote device can be affected, while some links may be inactive by design. The complete event and peer topology distinguish these cases.
What you'll observe
- Collective or transfer behavior changes while NVLink state or error counters differ from the expected topology.
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| nvidia-smi nvlink reports an inactive link or changing error counters | The link evidence must be correlated with expected topology and both endpoints. Xid 74 can be a consequence of the remote device failing rather than a defective local GPU. |
Which systems are affected
- NVIDIA platforms with NVLink connections or topology checks
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Compare nvidia-smi topology and link status with the expected machine configuration.
- ✓Correlate local and remote events, accounting for clock skew.
- ✓Check affected collective completion, output correctness and recurrence after approved recovery.
Root cause
- The link evidence must be correlated with expected topology and both endpoints. Xid 74 can be a consequence of the remote device failing rather than a defective local GPU.
The fix and how to prevent it
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Why the recommended fix works
Identifying the initiating endpoint prevents a downstream reporter being mistaken for the defective component.
Code examples
nvidia-smi topo -m
nvidia-smi nvlink -sAdapt the snippet to your framework. The same pattern holds for PyTorch Lightning, Hugging Face Trainer, DeepSpeed, Megatron-LM, and vLLM training wrappers. Where the wrapper exposes a config flag (for examplelr_scheduler_type in Trainer), prefer the flag over the imperative API to keep the schedule declarative and reproducible.
Best practices by model family
| Model / Stack | Recommendation | Notes |
|---|---|---|
| Small CNN / MLP | Recommended | Stabilises early-gradient noise even for tiny models. |
| Transformer (ViT/BERT) | Required | Attention stacks amplify gradient instability without active mitigation. |
| LLM (Llama / Qwen / GPT) | Required | At scale, every failure compounds across distributed collectives. |
| Diffusion / Stable Diffusion | Recommended | U-Net + cross-attention paths benefit from the same hardening. |
With the fix vs without the fix
| Dimension | With the fix | Without the fix |
|---|---|---|
| Symptom window | Stable from the first step | Visible within tens to hundreds of steps |
| Final metrics | Reproducible optima | Plateau or divergence below the baseline |
| Operational risk | Bounded by the prevention checklist | Compounds across folds / reruns |
| Prod recommendation | Ship | Block until the fix is in place |
Diagnostic note
“This is an evidence method, not a claim of tested physical NVLink recovery on the single-GPU replay machine.”
Diagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionRelated failures to investigate next
Frequently asked questions
Questions engineers and on-call staff commonly ask about this failure.
Is every inactive link a fault?
Does Xid 74 prove the reporting GPU is bad?
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.