GPU Thermal Slowdown or Power Brake? Check the Reason
Distinguish SW thermal slowdown, HW thermal slowdown and external power brake before changing cooling, power settings or the workload.
SW Thermal Slowdown indicates temperature-driven clock capping. HW Slowdown can include hardware thermal slowdown or an external power brake. Inspect the specific event reason and temperature over time; the combined label does not prove that every slowdown is overheating.
- Root cause
- The active clock event reason distinguishes thermal capping from an external power brake. Temperature, timing and system events are needed to explain the observed slowdown.
- Recommended fix
- Capture the full nvidia-smi query, active event reason, temperature and workload timing before changing settings.
- How Denpex helps
- Denpex investigates GPU Thermal Slowdown or Power Brake? Check the Reason using the evidence you provide or your connected workload collects. Earlier rank, host or application evidence is needed to distinguish an initiating failure from a downstream report.
What this failure is
NVIDIA clock event reasons describe why clocks are below requested values. Thermal and power-brake reasons need different evidence and operator actions.
Is this what broke your run? Paste your log.
You're reading about GPU Thermal Slowdown or Power Brake? Check the Reason. Paste your own traceback and relevant evidence for an investigation of your workload, with a next action or a specific missing fact. A reference entry does not establish your cause. No account or card for the free diagnosis. Review data handling before submitting sensitive logs.
Before uploading, review cloud data handling and local options.
Why it happens (the mechanism)
Thermal capping concerns temperature limits. An external power brake can concern system power delivery. Distinguish the active reason instead of assigning a single cooling cause to all clock reductions.
What you'll observe
- GPU clocks or workload throughput drop while clock event reasons change.
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| SW Thermal Slowdown | The active clock event reason distinguishes thermal capping from an external power brake. Temperature, timing and system events are needed to explain the observed slowdown. |
| HW Thermal Slowdown | The active clock event reason distinguishes thermal capping from an external power brake. Temperature, timing and system events are needed to explain the observed slowdown. |
| HW Power Brake | The active clock event reason distinguishes thermal capping from an external power brake. Temperature, timing and system events are needed to explain the observed slowdown. |
Which systems are affected
- NVIDIA GPUs and supported nvidia-smi clock event telemetry
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Read the specific active clock event reason for the installed GPU and driver.
- ✓Correlate reason changes with temperature and system events.
- ✓Verify intended workload results and throughput after the approved change.
Root cause
- The active clock event reason distinguishes thermal capping from an external power brake. Temperature, timing and system events are needed to explain the observed slowdown.
The fix and how to prevent it
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Why the recommended fix works
A targeted operator action addresses the documented thermal or power condition without hiding it behind an unsupported cap.
Code examples
nvidia-smi -q
# Read temperature and clock event reasons supported by this GPU/driver.Adapt the snippet to your framework. The same pattern holds for PyTorch Lightning, Hugging Face Trainer, DeepSpeed, Megatron-LM, and vLLM training wrappers. Where the wrapper exposes a config flag (for examplelr_scheduler_type in Trainer), prefer the flag over the imperative API to keep the schedule declarative and reproducible.
Best practices by model family
| Model / Stack | Recommendation | Notes |
|---|---|---|
| Small CNN / MLP | Recommended | Stabilises early-gradient noise even for tiny models. |
| Transformer (ViT/BERT) | Required | Attention stacks amplify gradient instability without active mitigation. |
| LLM (Llama / Qwen / GPT) | Required | At scale, every failure compounds across distributed collectives. |
| Diffusion / Stable Diffusion | Recommended | U-Net + cross-attention paths benefit from the same hardening. |
With the fix vs without the fix
| Dimension | With the fix | Without the fix |
|---|---|---|
| Symptom window | Stable from the first step | Visible within tens to hundreds of steps |
| Final metrics | Reproducible optima | Plateau or divergence below the baseline |
| Operational risk | Bounded by the prevention checklist | Compounds across folds / reruns |
| Prod recommendation | Ship | Block until the fix is in place |
Diagnostic note
“No physical cooling or power-delivery recovery experiment is claimed by this reference.”
Diagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionRelated failures to investigate next
Frequently asked questions
Questions engineers and on-call staff commonly ask about this failure.
Is HW Slowdown always overheating?
Should I set the power limit to 400 W?
References
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.