FSDP2 raises a KeyError on a tied weight while sharding the model
A language model whose output projection shares its tensor with the input embedding fails to shard. When fully_shard rebuilds parameters as DTensors, the shared tensor appears once rather than twice, the mapping from old parameters to new ones loses the tied name, and preparation dies on a KeyError naming a weight that plainly exists in the model.
The weight is not missing, it is an alias. The output projection shares its tensor with the embedding, so the shard mapping records it once. Disable weight tying in the config, or untie before sharding and re-tie after.
- Symptom
KeyError: 'lm_head.weight'- Root cause
- Weight tying means two parameter names refer to one tensor object. Code that walks named parameters therefore sees the tensor twice under two names, while code that builds a set or a dict keyed by the tensor sees it once. The mapping from pre-shard to post-shard parameters is built exactly that way, so the second name has no entry and the lookup raises.
- Recommended fix
- Disable weight tying in the model configuration and materialise a genuine tensor for the output projection before sharding, which is what the warning above the traceback is telling you to do.
- How Denpex helps
- Denpex matches FSDP2 raises a KeyError on a tied weight while sharding the model across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
What this failure is
A preparation-time failure in FSDP2 where two parameter names alias one tied tensor, so the mapping rebuilt during sharding records the tensor once and a lookup by the second name raises a KeyError.
Is this what broke your run? Paste your log.
You're reading about FSDP2 raises a KeyError on a tied weight while sharding the model. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.
Why it happens (the mechanism)
Sharding replaces every parameter with a new object and needs to know which new object corresponds to which old one. That correspondence is built from tensor identity, and identity is precisely what tying collapses. Two names go in, one entry comes out, and the name that lost the race is the one the error reports.
What you'll observe
- The named weight is visibly present in the model, so the error reads as impossible
- The failure happens during preparation, before any training step
- The same model trains without complaint under the previous FSDP implementation
- A warning about weight tying appears above the traceback and is easy to read past
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| KeyError: 'lm_head.weight' raised while preparing the model | Weight tying means two parameter names refer to one tensor object. Code that walks named parameters therefore sees the tensor twice under two names, while code that builds a set or a dict keyed by the tensor sees it once. The mapping from pre-shard to post-shard parameters is built exactly that way, so the second name has no entry and the lookup raises. |
| Traceback through accelerator.prepare into _prepare_fsdp2, at a mapping built from old and new named parameters | FSDP2 is more exposed to this than its predecessor because it replaces each parameter with a DTensor carrying its own placement. A tensor shared across two shard groups has no single well-defined placement, so tying is not a cosmetic detail the sharding layer can ignore. |
| A preceding warning suggesting the config be updated with tie_word_embeddings set to false | The error names the weight rather than the tying, which sends people to look for a missing parameter. Nothing is missing: the name is a second alias for a tensor the mapping already consumed under its first alias. |
| KeyError reporting that a parameter in the optimizer could not be switched to its sharded version | Weight tying means two parameter names refer to one tensor object. Code that walks named parameters therefore sees the tensor twice under two names, while code that builds a set or a dict keyed by the tensor sees it once. The mapping from pre-shard to post-shard parameters is built exactly that way, so the second name has no entry and the lookup raises. |
Which systems are affected
- FSDP2 and fully_shard, which represent sharded parameters as DTensors
- Causal language models that tie the output projection to the input embedding, which is most of them
- Setups where the tied tensors would land in different shard groups, such as the embedding at the root and the head below it
- Optimizer state handling, where the same missing key surfaces on save rather than on prepare
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Compare the identity of the output projection tensor with the input embedding tensor. If they are the same object, tying is in play and explains the missing name.
- ✓Read the lines immediately above the traceback for a warning about tied weights; it names the configuration change that resolves it.
- ✓Re-run with tying disabled in the configuration. Success confirms the alias rather than a genuinely absent parameter.
Root cause
- Weight tying means two parameter names refer to one tensor object. Code that walks named parameters therefore sees the tensor twice under two names, while code that builds a set or a dict keyed by the tensor sees it once. The mapping from pre-shard to post-shard parameters is built exactly that way, so the second name has no entry and the lookup raises.
- FSDP2 is more exposed to this than its predecessor because it replaces each parameter with a DTensor carrying its own placement. A tensor shared across two shard groups has no single well-defined placement, so tying is not a cosmetic detail the sharding layer can ignore.
- The error names the weight rather than the tying, which sends people to look for a missing parameter. Nothing is missing: the name is a second alias for a tensor the mapping already consumed under its first alias.
The fix and how to prevent it
Searchable error signature
KeyError: 'lm_head.weight'
File ".../accelerate/accelerator.py", in _prepare_fsdp2
mapping = {p: new_named_params[n] for n, p in old_named_params.items()}
WARNING: model has tied weights; consider setting tie_word_embeddings=False in the configUse this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Why the recommended fix works
Giving the output projection its own tensor makes the two names refer to two objects, so the mapping has an entry for each and the lookup succeeds. Untie-shard-retie achieves the same thing for a window long enough to build the mapping. Neither changes the mathematics of the model; they change whether the sharding layer can express it.
Best practices by model family
| Model / Stack | Recommendation | Notes |
|---|---|---|
| Causal LM with tied embeddings | Set tie_word_embeddings to false before sharding | Gives the head its own tensor so the mapping has two entries. |
| Tying must be preserved | Untie, shard, then restore the relationship | The ambiguity only exists while the mapping is built. |
| Failure appears on save | Same fix; the optimizer path shares the mapping | The missing name breaks state handling identically. |
| Loading a state dict without the tied name | Load non-strictly and copy the embedding into the slot | The value is available under the other alias. |
With the fix vs without the fix
| Dimension | With the fix | Without the fix |
|---|---|---|
| What the KeyError means | A name that aliases an already-mapped tensor | Read as a parameter missing from the model |
| Why FSDP2 specifically | Each parameter becomes a DTensor with its own placement | Assumed to behave like the previous implementation |
| The warning above the traceback | Names the configuration change that fixes it | Skipped as routine noise |
Diagnostic note
“The reflex on seeing a KeyError for a weight is to check whether the checkpoint is missing it, and that is always a dead end here. The tell is that the name in the error is the output projection specifically, which is the tied one in almost every causal language model. If you see any tied-weight name in a sharding traceback, go straight to the tying configuration rather than to the checkpoint.”
Visual fingerprint
before sharding
model.embed_tokens.weight -> tensor A
lm_head.weight -> tensor A (same object)
mapping keyed by tensor identity
tensor A -> new DTensor (one entry, not two)
lookup by name 'lm_head.weight' -> KeyErrorDiagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionRelated failures to investigate next
Frequently asked questions
Questions engineers and on-call staff commonly ask about this failure.
The parameter is clearly in my model. Why a KeyError?
It worked under the older FSDP.
Will untying change my model?
Mine fails on save, not on prepare.
References
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.