Skip to content

FSDP2 raises a KeyError on a tied weight while sharding the model

A language model whose output projection shares its tensor with the input embedding fails to shard. When fully_shard rebuilds parameters as DTensors, the shared tensor appears once rather than twice, the mapping from old parameters to new ones loses the tied name, and preparation dies on a KeyError naming a weight that plainly exists in the model.

Quick answer

The weight is not missing, it is an alias. The output projection shares its tensor with the embedding, so the shard mapping records it once. Disable weight tying in the config, or untie before sharding and re-tie after.

Symptom
KeyError: 'lm_head.weight'
Root cause
Weight tying means two parameter names refer to one tensor object. Code that walks named parameters therefore sees the tensor twice under two names, while code that builds a set or a dict keyed by the tensor sees it once. The mapping from pre-shard to post-shard parameters is built exactly that way, so the second name has no entry and the lookup raises.
Recommended fix
Disable weight tying in the model configuration and materialise a genuine tensor for the output projection before sharding, which is what the warning above the traceback is telling you to do.
How Denpex helps
Denpex matches FSDP2 raises a KeyError on a tied weight while sharding the model across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
Distributed Training#fsdp#fsdp2#fully_shard#tied-weights#lm_head#keyerror

What this failure is

A preparation-time failure in FSDP2 where two parameter names alias one tied tensor, so the mapping rebuilt during sharding records the tensor once and a lookup by the second name raises a KeyError.

Live diagnosis, no signup

Is this what broke your run? Paste your log.

You're reading about FSDP2 raises a KeyError on a tied weight while sharding the model. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.

training_logs.txt
No log to hand? Try one:

3 free diagnoses/day

Why it happens (the mechanism)

Sharding replaces every parameter with a new object and needs to know which new object corresponds to which old one. That correspondence is built from tensor identity, and identity is precisely what tying collapses. Two names go in, one entry comes out, and the name that lost the race is the one the error reports.

What you'll observe

  • The named weight is visibly present in the model, so the error reads as impossible
  • The failure happens during preparation, before any training step
  • The same model trains without complaint under the previous FSDP implementation
  • A warning about weight tying appears above the traceback and is easy to read past

Common symptoms and what they mean

SymptomWhy it happens
KeyError: 'lm_head.weight' raised while preparing the modelWeight tying means two parameter names refer to one tensor object. Code that walks named parameters therefore sees the tensor twice under two names, while code that builds a set or a dict keyed by the tensor sees it once. The mapping from pre-shard to post-shard parameters is built exactly that way, so the second name has no entry and the lookup raises.
Traceback through accelerator.prepare into _prepare_fsdp2, at a mapping built from old and new named parametersFSDP2 is more exposed to this than its predecessor because it replaces each parameter with a DTensor carrying its own placement. A tensor shared across two shard groups has no single well-defined placement, so tying is not a cosmetic detail the sharding layer can ignore.
A preceding warning suggesting the config be updated with tie_word_embeddings set to falseThe error names the weight rather than the tying, which sends people to look for a missing parameter. Nothing is missing: the name is a second alias for a tensor the mapping already consumed under its first alias.
KeyError reporting that a parameter in the optimizer could not be switched to its sharded versionWeight tying means two parameter names refer to one tensor object. Code that walks named parameters therefore sees the tensor twice under two names, while code that builds a set or a dict keyed by the tensor sees it once. The mapping from pre-shard to post-shard parameters is built exactly that way, so the second name has no entry and the lookup raises.

Which systems are affected

  • FSDP2 and fully_shard, which represent sharded parameters as DTensors
  • Causal language models that tie the output projection to the input embedding, which is most of them
  • Setups where the tied tensors would land in different shard groups, such as the embedding at the root and the head below it
  • Optimizer state handling, where the same missing key surfaces on save rather than on prepare

How to confirm this is the problem

Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.

  • ✓Compare the identity of the output projection tensor with the input embedding tensor. If they are the same object, tying is in play and explains the missing name.
  • ✓Read the lines immediately above the traceback for a warning about tied weights; it names the configuration change that resolves it.
  • ✓Re-run with tying disabled in the configuration. Success confirms the alias rather than a genuinely absent parameter.

Root cause

  • Weight tying means two parameter names refer to one tensor object. Code that walks named parameters therefore sees the tensor twice under two names, while code that builds a set or a dict keyed by the tensor sees it once. The mapping from pre-shard to post-shard parameters is built exactly that way, so the second name has no entry and the lookup raises.
  • FSDP2 is more exposed to this than its predecessor because it replaces each parameter with a DTensor carrying its own placement. A tensor shared across two shard groups has no single well-defined placement, so tying is not a cosmetic detail the sharding layer can ignore.
  • The error names the weight rather than the tying, which sends people to look for a missing parameter. Nothing is missing: the name is a second alias for a tensor the mapping already consumed under its first alias.

The fix and how to prevent it

Searchable error signature

search key
KeyError: 'lm_head.weight'
  File ".../accelerate/accelerator.py", in _prepare_fsdp2
    mapping = {p: new_named_params[n] for n, p in old_named_params.items()}
WARNING: model has tied weights; consider setting tie_word_embeddings=False in the config

Use this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.

The fix and the prevention pattern

The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.

Sign up free. Unlock the full analysis

No credit card. Daily allowance follows verified trust tier. Instant access.

Why the recommended fix works

Giving the output projection its own tensor makes the two names refer to two objects, so the mapping has an entry for each and the lookup succeeds. Untie-shard-retie achieves the same thing for a window long enough to build the mapping. Neither changes the mathematics of the model; they change whether the sharding layer can express it.

Best practices by model family

Model / StackRecommendationNotes
Causal LM with tied embeddingsSet tie_word_embeddings to false before shardingGives the head its own tensor so the mapping has two entries.
Tying must be preservedUntie, shard, then restore the relationshipThe ambiguity only exists while the mapping is built.
Failure appears on saveSame fix; the optimizer path shares the mappingThe missing name breaks state handling identically.
Loading a state dict without the tied nameLoad non-strictly and copy the embedding into the slotThe value is available under the other alias.

With the fix vs without the fix

DimensionWith the fixWithout the fix
What the KeyError meansA name that aliases an already-mapped tensorRead as a parameter missing from the model
Why FSDP2 specificallyEach parameter becomes a DTensor with its own placementAssumed to behave like the previous implementation
The warning above the tracebackNames the configuration change that fixes itSkipped as routine noise

Diagnostic note

“The reflex on seeing a KeyError for a weight is to check whether the checkpoint is missing it, and that is always a dead end here. The tell is that the name in the error is the output projection specifically, which is the tied one in almost every causal language model. If you see any tied-weight name in a sharding traceback, go straight to the tying configuration rather than to the checkpoint.”

Visual fingerprint

Two names, one tensor, one mapping entry
  before sharding
    model.embed_tokens.weight  ->  tensor A
    lm_head.weight             ->  tensor A     (same object)

  mapping keyed by tensor identity
    tensor A -> new DTensor        (one entry, not two)

  lookup by name 'lm_head.weight' -> KeyError
Both names point at one tensor, so a mapping keyed by tensor identity records a single entry. Looking that mapping up by the second name finds nothing, which is the KeyError, even though the parameter is present in the model.

Diagnose this failure in VS Code

Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.

Install the free VS Code extension

Frequently asked questions

Questions engineers and on-call staff commonly ask about this failure.

The parameter is clearly in my model. Why a KeyError?
It is present under two names that share one tensor. The shard mapping is keyed by tensor identity, so it holds one entry and the second name finds nothing.
It worked under the older FSDP.
FSDP2 gives each parameter its own DTensor placement, so a tensor shared across shard groups has no single valid placement. The older implementation did not need to resolve that.
Will untying change my model?
It stops the two tensors being one object. If you need the weights to stay equal, untie only around sharding and restore the relationship afterwards.
Mine fails on save, not on prepare.
Same cause. The optimizer state path uses the same mapping, and the missing alias breaks it there too.

Don't just read the fix, diagnose your run

The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.