FSDP fails when composed with CPU parameter offloading
Sharding and CPU offloading both move parameters, and each assumes it is the component deciding where a parameter lives. Enabling them together produces a failure during setup or the first step, in a stack that works correctly with either one alone.
Two memory managers are claiming the same parameters. Test each alone to confirm, then recover the memory from the activation side or from wider sharding rather than from offloading.
- Symptom
File ".../megatron/core/optimizer/cpu_offloading/hybrid_optimizer.py", line 188, in _init_sub_optimizers- Root cause
- The concrete mechanism in the Megatron implementation is worth stating exactly, because it explains why the error names the optimizer rather than the sharding. The offload path builds its host-side copy by calling detach, clone, cpu and pin_memory on each parameter in turn. Under sharding those parameters are no longer plain tensors but DTensors, so every one of those calls goes through a dispatcher that needs a sharding rule for the operator.
- Recommended fix
- Establish which of the two is unsupported in combination for the exact versions in use, by testing each alone before concluding anything about the pair.
- How Denpex helps
- Denpex matches FSDP fails when composed with CPU parameter offloading across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
What this failure is
A composition failure in which parameter sharding and CPU offloading each claim authority over where a parameter resides, producing a setup-time or first-step error in a configuration where either feature alone succeeds.
Is this what broke your run? Paste your log.
You're reading about FSDP fails when composed with CPU parameter offloading. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.
Why it happens (the mechanism)
Both features exist to reduce resident parameter memory and both work by controlling placement. Sharding splits a parameter and records where each piece lives; offloading moves it to host memory and records that instead. Neither is written to defer to the other, so the composition has no agreed owner for a tensor, and the first thing that resolves placement finds a contradiction.
What you'll observe
- Both features are documented and each works by itself
- The failure appears at setup or on the first step rather than under memory pressure
- Offloading was enabled precisely because memory was already tight, so simply removing it is not free
- The error rarely names offloading, so the interaction is not obvious from the message
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| The traceback ending inside HybridDeviceOptimizer while it builds its sub-optimizer parameter groups, not inside the sharding layer | The concrete mechanism in the Megatron implementation is worth stating exactly, because it explains why the error names the optimizer rather than the sharding. The offload path builds its host-side copy by calling detach, clone, cpu and pin_memory on each parameter in turn. Under sharding those parameters are no longer plain tensors but DTensors, so every one of those calls goes through a dispatcher that needs a sharding rule for the operator. Pinning host memory has no meaningful sharded semantics, the rule is absent, and the composition dies in the optimizer constructor. |
| A call to param.detach().clone().cpu().pin_memory() applied to a parameter that sharding has already replaced with a DTensor | Sharding divides a parameter across ranks and keeps authoritative placement metadata for each piece. Offloading relocates parameters to host memory and restores them on demand, keeping its own view of where each one is. Both are correct in isolation and each expects to be the authority, so composing them leaves two components disagreeing about a single tensor. |
| Frames passing through the DTensor dispatcher and its sharding propagator on the way to the failure | The disagreement shows up early because placement is resolved during setup and on the first use of a parameter, not when memory runs short. That is why the failure looks unrelated to the memory pressure that motivated enabling offloading. |
| The failure raised from get_megatron_optimizer during setup_model_and_optimizer, before the first training step | Support for the combination is a property of a specific implementation and version rather than of the idea, so a stack that documents both features separately may not support them together, and a version that works is not evidence a neighbouring one does. |
| Setup or the first training step failing with both sharding and CPU offloading enabled | The concrete mechanism in the Megatron implementation is worth stating exactly, because it explains why the error names the optimizer rather than the sharding. The offload path builds its host-side copy by calling detach, clone, cpu and pin_memory on each parameter in turn. Under sharding those parameters are no longer plain tensors but DTensors, so every one of those calls goes through a dispatcher that needs a sharding rule for the operator. Pinning host memory has no meaningful sharded semantics, the rule is absent, and the composition dies in the optimizer constructor. |
| A device or placement error naming a parameter that offloading has moved to host memory | Sharding divides a parameter across ranks and keeps authoritative placement metadata for each piece. Offloading relocates parameters to host memory and restores them on demand, keeping its own view of where each one is. Both are correct in isolation and each expects to be the authority, so composing them leaves two components disagreeing about a single tensor. |
Which systems are affected
- Megatron-style FSDP implementations combined with CPU offloading
- Stacks where offloading is configured by a launcher and sharding by the training framework, so neither knows about the other
- Configurations carried over from a single-device memory-saving setup into a sharded one
- Any setup where a parameter's owner is ambiguous between two memory managers
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Read which frame raises. Landing in the optimizer's offload helper rather than in the sharding layer tells you offloading is the caller and sharding is only the thing it tripped over.
- ✓Look for a DTensor dispatch frame beneath the offload call. Its presence means the parameter was already sharded when the host copy was attempted.
- ✓Disable CPU offloading and re-run with sharding unchanged. Success isolates the interaction rather than either feature.
- ✓Disable sharding and re-run with offloading unchanged. Success on both single-feature runs confirms the composition is the problem.
- ✓Check whether the failure occurs at setup or on the first step rather than under sustained memory pressure, which distinguishes a placement disagreement from genuine exhaustion.
Root cause
- The concrete mechanism in the Megatron implementation is worth stating exactly, because it explains why the error names the optimizer rather than the sharding. The offload path builds its host-side copy by calling detach, clone, cpu and pin_memory on each parameter in turn. Under sharding those parameters are no longer plain tensors but DTensors, so every one of those calls goes through a dispatcher that needs a sharding rule for the operator. Pinning host memory has no meaningful sharded semantics, the rule is absent, and the composition dies in the optimizer constructor.
- Sharding divides a parameter across ranks and keeps authoritative placement metadata for each piece. Offloading relocates parameters to host memory and restores them on demand, keeping its own view of where each one is. Both are correct in isolation and each expects to be the authority, so composing them leaves two components disagreeing about a single tensor.
- The disagreement shows up early because placement is resolved during setup and on the first use of a parameter, not when memory runs short. That is why the failure looks unrelated to the memory pressure that motivated enabling offloading.
- Support for the combination is a property of a specific implementation and version rather than of the idea, so a stack that documents both features separately may not support them together, and a version that works is not evidence a neighbouring one does.
The fix and how to prevent it
Searchable error signature
File ".../megatron/core/optimizer/cpu_offloading/hybrid_optimizer.py", line 188, in _init_sub_optimizers
) = self._get_sub_optimizer_param_groups(self.offload_fraction)
param = param.detach().clone().cpu().pin_memory()
File ".../torch/distributed/tensor/_dispatch.py", line 160, in dispatch
self.sharding_propagator.propagate(op_info)
Megatron FSDP does not work with CPU_OFFLOADING
RuntimeError raised during setup with both parameter sharding and CPU offloading enabledUse this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Why the recommended fix works
Removing one claimant leaves a single authority and the contradiction disappears. Recovering the memory from activations works because activation memory is managed by a different mechanism entirely, so it does not compete for ownership of parameters. Widening sharding reduces the same footprint offloading was targeting, without a second manager.
Best practices by model family
| Model / Stack | Recommendation | Notes |
|---|---|---|
| Sharding must stay | Widen sharding or shard optimizer state further | Reduces the same footprint without a second memory manager. |
| Memory pressure is from activations | Activation checkpointing or a smaller microbatch | Activation memory does not compete for parameter ownership. |
| Offloading is essential | Pin a version where the pair is supported | Support is a property of the implementation and moves between releases. |
| Diagnosing | Test each feature alone first | Two successful single-feature runs prove it is the composition. |
With the fix vs without the fix
| Dimension | With the fix | Without the fix |
|---|---|---|
| What each feature controls | Both control parameter placement | Assumed to address different resources |
| When the failure appears | At setup or the first step | Expected under sustained memory pressure |
| Whether documentation implies support | Documented separately, not necessarily together | Read as both being supported at once |
Diagnostic note
“The instinct on hitting this is to look for a memory bug, because offloading was enabled under memory pressure and the failure arrives soon after. The timing is the clue that it is not: a genuine exhaustion appears when the workload grows, while a placement disagreement appears at setup or on the very first step regardless of size. Note which, before spending time on allocator settings.”
Visual fingerprint
sharding : parameter split across ranks, placement metadata per shard
offloading : parameter relocated to host memory, restored on demand
both enabled -> who owns the placement of this tensor?
|
resolved at setup / first use -> contradictionDiagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionFrequently asked questions
Questions engineers and on-call staff commonly ask about this failure.
Both features are documented. Why can I not use both?
Is this an out-of-memory problem?
I enabled offloading because I was out of memory.
How do I confirm it is the combination?
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.