MXFP4 expert weights lose their scale attribute when a mixture-of-experts model loads under FSDP2
A quantised mixture-of-experts checkpoint loads on a single device and fails under sharded loading. Materialising a module from the placeholder device reads each parameter by name, and the quantised expert block does not expose the scale attribute the checkpoint expects, so loading stops on a missing attribute with a helpful suggestion naming a different one.
The quantised expert module and the checkpoint disagree about what the scale tensors are called, and sharded loading resolves names against the live module rather than into a state dictionary. Load first and shard afterwards, and ignore the attribute the error suggests.
- Symptom
AttributeError: 'Mxfp4GptOssExperts' object has no attribute 'down_proj_scales'. Did you mean: 'down_proj_bias'?- Root cause
- A quantised expert block does not hold the same tensors as the unquantised one. The packed weights and their scales may be registered under different names, fused into one buffer, or reconstructed on first use, and which of those a given implementation chooses is an internal detail that the checkpoint's key layout does not have to agree with. Sharded loading makes that disagreement fatal.
- Recommended fix
- Load the checkpoint without sharding first and shard the materialised model afterwards, which avoids the name-by-name materialisation path that the quantised module does not satisfy.
- How Denpex helps
- Denpex matches MXFP4 expert weights lose their scale attribute when a mixture-of-experts model loads under FSDP2 across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
What this failure is
A load-time failure in which sharded materialisation resolves a checkpoint key against a live quantised module by attribute name, and the quantised expert implementation does not expose the scale tensor under the name the checkpoint uses.
Is this what broke your run? Paste your log.
You're reading about MXFP4 expert weights lose their scale attribute when a mixture-of-experts model loads under FSDP2. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.
Why it happens (the mechanism)
Sharded construction is designed to avoid ever holding the whole model, so it builds an empty skeleton and fills each parameter in as its key is read. That requires the module to already expose every name the checkpoint will mention. A quantised implementation is free to store packed weights and scales however it likes, and when its choice differs from the checkpoint's key layout, the discrepancy surfaces here rather than in a conventional load that assigns into a dictionary first.
What you'll observe
- The identical checkpoint loads correctly without sharding, so the file is plainly fine
- The error names a missing attribute and suggests a near neighbour, which invites editing the name rather than understanding it
- It fails during loading, before any training step, so no gradient or collective is involved
- Quantised experts and sharded loading are each supported, and only their combination fails
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| AttributeError naming a quantised expert module and an absent scale attribute, with a suggestion pointing at the corresponding bias | A quantised expert block does not hold the same tensors as the unquantised one. The packed weights and their scales may be registered under different names, fused into one buffer, or reconstructed on first use, and which of those a given implementation chooses is an internal detail that the checkpoint's key layout does not have to agree with. |
| The failure raised while a state dictionary is loaded into a model still on the placeholder device | Sharded loading makes that disagreement fatal. Building the model without data and then materialising parameter by parameter requires the attribute to exist on the module at the moment its name comes up, whereas a conventional load can assign into a state dictionary and let the module sort itself out afterwards. |
| Checkpoint shard loading reaching completion on some ranks and stopping partway on another | The suggestion in the error is misleading and worth ignoring. It is a spelling hint generated from the attributes that do exist, and the neighbour it proposes is a different tensor entirely; acting on it would load a bias where a scale belongs. |
| A launcher reporting a non-zero exit code for local rank zero with no collective error beneath it | A quantised expert block does not hold the same tensors as the unquantised one. The packed weights and their scales may be registered under different names, fused into one buffer, or reconstructed on first use, and which of those a given implementation chooses is an internal detail that the checkpoint's key layout does not have to agree with. |
Which systems are affected
- Mixture-of-experts models whose expert weights are stored in a block-scaled four-bit format
- Fully sharded loading, which builds the model without data and fills parameters in by name
- Quantised module implementations that register some tensors as parameters and derive or fuse others
- Any loader that pairs a checkpoint key with an attribute lookup on the live module
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Load the same checkpoint on one device with no sharding. Succeeding there confirms the sharded materialisation path rather than the checkpoint.
- ✓List the expert module's parameters and buffers and compare the names against the checkpoint keys. The absent name is the disagreement, stated exactly.
- ✓Check the quantisation library version against the one that produced the checkpoint, since these attribute names change with the scheme.
Root cause
- A quantised expert block does not hold the same tensors as the unquantised one. The packed weights and their scales may be registered under different names, fused into one buffer, or reconstructed on first use, and which of those a given implementation chooses is an internal detail that the checkpoint's key layout does not have to agree with.
- Sharded loading makes that disagreement fatal. Building the model without data and then materialising parameter by parameter requires the attribute to exist on the module at the moment its name comes up, whereas a conventional load can assign into a state dictionary and let the module sort itself out afterwards.
- The suggestion in the error is misleading and worth ignoring. It is a spelling hint generated from the attributes that do exist, and the neighbour it proposes is a different tensor entirely; acting on it would load a bias where a scale belongs.
The fix and how to prevent it
Searchable error signature
AttributeError: 'Mxfp4GptOssExperts' object has no attribute 'down_proj_scales'. Did you mean: 'down_proj_bias'?
File ".../transformers/modeling_utils.py", in _load_state_dict_into_meta_model
value = getattr(module, param_type)
Loading checkpoint shards: 0%| | 0/3 [00:00<?, ?it/s]Use this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Why the recommended fix works
Loading before sharding restores the ordinary path, where the checkpoint is read into a fully constructed model and the quantised module can reconcile its own tensors before anything is sharded. Aligning the quantisation library with the one that wrote the checkpoint removes the disagreement at its source. Both mean the name being looked up is one the module actually has.
Best practices by model family
| Model / Stack | Recommendation | Notes |
|---|---|---|
| Quantised MoE under sharding | Load first, shard afterwards | Avoids name-by-name materialisation. |
| Memory permits | Load the unquantised form | The mismatch is in the quantised tensor layout. |
| Checkpoint from elsewhere | Match the quantisation library version | Attribute names are part of the format. |
| Diagnosing | Diff module attributes against checkpoint keys | States the disagreement exactly. |
With the fix vs without the fix
| Dimension | With the fix | Without the fix |
|---|---|---|
| Why single-device loading works | Assigns into a state dictionary first | Assumed to be the same code path |
| What the Did you mean hint means | A spelling neighbour, nothing more | Read as the correct attribute |
| Where the disagreement lives | Between module attributes and checkpoint keys | Blamed on the sharding layer |
Diagnostic note
“The suggested attribute in that error has cost people real time. A missing-attribute message with a Did you mean hint is generated by scanning what does exist for a similar spelling, and similarity of spelling says nothing about similarity of meaning. A bias and a scale differ by one word in the name and by everything in what they contain. Treat the hint as noise whenever the attributes are tensors rather than methods.”
Visual fingerprint
conventional load
build full model -> read checkpoint into state dict -> module reconciles ✓
sharded materialisation
build skeleton on meta
for each key: getattr(module, name) <- name must ALREADY exist
'down_proj_scales' absent -> AttributeErrorDiagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionFrequently asked questions
Questions engineers and on-call staff commonly ask about this failure.
Should I use the attribute it suggests?
Why does it load fine on one GPU?
Is the checkpoint broken?
Can I keep sharding?
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.