Weights-only load failed with an UnpicklingError when you resume an older checkpoint
A checkpoint that loaded without complaint for months starts failing after a framework upgrade. Nothing about the file changed. The loader now refuses by default to reconstruct any Python object that is not a plain tensor, and a training checkpoint carries several, the optimizer class among them, so the load stops on the first one it meets.
The checkpoint is fine; the loader got stricter. It refuses to rebuild Python objects like the optimizer class. Allowlist the exact class the error names, ideally through the scoped context manager, and expect one or two more behind it.
- Symptom
_pickle.UnpicklingError: Weights only load failed. This file can still be loaded, to do so you have two options- Root cause
- Reconstructing a pickle can execute arbitrary code, because the format stores instructions for rebuilding objects rather than only their data. Restricting the loader to tensors and primitives by default closes that hole, and it is the right default for a file arriving from anywhere but your own cluster. A training checkpoint is not only tensors.
- Recommended fix
- Allowlist the specific classes the checkpoint legitimately contains, using the serialization allowlist the error message names. This keeps the protection for everything else and is the fix that is safe to leave in place.
- How Denpex helps
- Denpex matches Weights-only load failed with an UnpicklingError when you resume an older checkpoint across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
What this failure is
A resume-time failure in which a checkpoint written under a permissive deserialization default is read under a restrictive one, so the loader refuses to reconstruct the non-tensor objects the checkpoint legitimately contains.
Is this what broke your run? Paste your log.
You're reading about Weights-only load failed with an UnpicklingError when you resume an older checkpoint. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.
Why it happens (the mechanism)
Deserialization of a pickle is code execution, so the safe default is to rebuild nothing but tensors. Training checkpoints have always stored more than tensors, because resuming needs the optimizer, the scheduler and the run configuration. Those two facts were compatible only while the default was permissive; once it flipped, every previously written checkpoint became a file the reader declines to fully trust.
What you'll observe
- The file is intact and readable, so every corruption check comes back clean
- The same file loaded successfully under the previous framework release
- The message offers a way to disable the restriction, which is the fastest fix and the wrong default
- It surfaces at resume, typically hours into a queue slot, on every rank at once
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| _pickle.UnpicklingError: Weights only load failed, raised from inside the checkpoint loader | Reconstructing a pickle can execute arbitrary code, because the format stores instructions for rebuilding objects rather than only their data. Restricting the loader to tensors and primitives by default closes that hole, and it is the right default for a file arriving from anywhere but your own cluster. |
| WeightsUnpickler error: Unsupported global, naming a class such as an optimizer that was not an allowed global by default | A training checkpoint is not only tensors. It holds the optimizer, its parameter groups, the scheduler, the sampler position and the run's arguments, all of which are Python objects that the restricted loader will not rebuild. The first one encountered stops the load, so the reported class is whichever came first and not necessarily the only one. |
| A suggestion to call add_safe_globals or use the safe_globals context manager with the named class | Nothing is wrong with the checkpoint. It was written under a permissive default and is being read under a strict one, so the change of behaviour lives entirely in the reader. This is why file-level integrity checks all pass and why the same bytes still load in the older release. |
| Every rank raising the same error at the same point, because they are all reading the same common state | Reconstructing a pickle can execute arbitrary code, because the format stores instructions for rebuilding objects rather than only their data. Restricting the loader to tensors and primitives by default closes that hole, and it is the right default for a file arriving from anywhere but your own cluster. |
Which systems are affected
- Resuming any training checkpoint written before the loader's default changed
- Distributed checkpoint formats that store a common state dictionary alongside the sharded tensors
- Optimizer state, which pickles the optimizer class itself rather than only its numbers
- Learning-rate schedulers, dataloader samplers, and argument namespaces saved beside the weights
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Read the class named after Unsupported global. If it is an optimizer, a scheduler or an argument container, this is the strict-loader change and not damage to the file.
- ✓Load the same file with the restriction disabled once, in a scratch process. Success proves the bytes are fine and the reader's default is what moved.
- ✓Check whether the checkpoint predates your current framework release. A file written under the permissive default is the whole precondition.
Root cause
- Reconstructing a pickle can execute arbitrary code, because the format stores instructions for rebuilding objects rather than only their data. Restricting the loader to tensors and primitives by default closes that hole, and it is the right default for a file arriving from anywhere but your own cluster.
- A training checkpoint is not only tensors. It holds the optimizer, its parameter groups, the scheduler, the sampler position and the run's arguments, all of which are Python objects that the restricted loader will not rebuild. The first one encountered stops the load, so the reported class is whichever came first and not necessarily the only one.
- Nothing is wrong with the checkpoint. It was written under a permissive default and is being read under a strict one, so the change of behaviour lives entirely in the reader. This is why file-level integrity checks all pass and why the same bytes still load in the older release.
The fix and how to prevent it
Searchable error signature
_pickle.UnpicklingError: Weights only load failed. This file can still be loaded, to do so you have two options
WeightsUnpickler error: Unsupported global: GLOBAL torch.optim.adamw.AdamW was not an allowed global by default
Please use `torch.serialization.add_safe_globals([torch.optim.adamw.AdamW])` to allowlist this global
File ".../megatron/training/checkpointing.py", in _load_base_checkpoint
state_dict = dist_checkpointing.load_common_state_dict(checkpoint_name)Use this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Why the recommended fix works
Allowlisting states, class by class, that you know what is in your own file. The loader keeps refusing everything you did not name, so the protection still applies to any checkpoint arriving from elsewhere. Disabling the restriction wholesale also loads the file, but it discards the protection for every object in it rather than for the one you inspected.
Best practices by model family
| Model / Stack | Recommendation | Notes |
|---|---|---|
| Your own cluster's checkpoint | Allowlist the named classes | Keeps the restriction for everything you did not name. |
| One-off load in a script | Scoped safe-globals context manager | The relaxation ends with the call. |
| Checkpoint from an external source | Allowlist only, never disable | This is exactly the case the default protects. |
| Fleet-wide upgrade | Re-save once in the strict format | Removes the need for an allowlist on future resumes. |
With the fix vs without the fix
| Dimension | With the fix | Without the fix |
|---|---|---|
| Is the checkpoint damaged | No, only read under a stricter default | Investigated as corruption |
| How many classes will it name | One at a time, in the order encountered | Assumed to be the only one |
| Where to relax the rule | At the single call site | As a process-wide setting |
Diagnostic note
“The message helpfully offers two routes and people take the first one, because it is a single argument and the job is already an hour into its allocation. That choice then propagates: it gets committed to the training script, and from there into every job on the cluster, including the ones loading checkpoints pulled from a model hub. Spend the extra minute on the allowlist. The class names are printed for you.”
Visual fingerprint
checkpoint
├── model tensors -> rebuilt under the strict default
├── optimizer state
│ └── class AdamW -> REFUSED: unsupported global
├── lr scheduler
└── run arguments
load stops at the FIRST refusal; the ones below are not yet reportedDiagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionFrequently asked questions
Questions engineers and on-call staff commonly ask about this failure.
Is my checkpoint corrupted?
Why not just disable the restriction?
I allowlisted the class and got another error.
Can I stop this happening on every resume?
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.