Skip to content

FSDP fails when composed with CPU parameter offloading

Sharding and CPU offloading both move parameters, and each assumes it is the component deciding where a parameter lives. Enabling them together produces a failure during setup or the first step, in a stack that works correctly with either one alone.

Quick answer

Two memory managers are claiming the same parameters. Test each alone to confirm, then recover the memory from the activation side or from wider sharding rather than from offloading.

Symptom
File ".../megatron/core/optimizer/cpu_offloading/hybrid_optimizer.py", line 188, in _init_sub_optimizers
Root cause
The concrete mechanism in the Megatron implementation is worth stating exactly, because it explains why the error names the optimizer rather than the sharding. The offload path builds its host-side copy by calling detach, clone, cpu and pin_memory on each parameter in turn. Under sharding those parameters are no longer plain tensors but DTensors, so every one of those calls goes through a dispatcher that needs a sharding rule for the operator.
Recommended fix
Establish which of the two is unsupported in combination for the exact versions in use, by testing each alone before concluding anything about the pair.
How Denpex helps
Denpex matches FSDP fails when composed with CPU parameter offloading across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
Distributed Training#fsdp#cpu-offload#offloading#composition#megatron#memory-saving

What this failure is

A composition failure in which parameter sharding and CPU offloading each claim authority over where a parameter resides, producing a setup-time or first-step error in a configuration where either feature alone succeeds.

Live diagnosis, no signup

Is this what broke your run? Paste your log.

You're reading about FSDP fails when composed with CPU parameter offloading. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.

training_logs.txt
No log to hand? Try one:

3 free diagnoses/day

Why it happens (the mechanism)

Both features exist to reduce resident parameter memory and both work by controlling placement. Sharding splits a parameter and records where each piece lives; offloading moves it to host memory and records that instead. Neither is written to defer to the other, so the composition has no agreed owner for a tensor, and the first thing that resolves placement finds a contradiction.

What you'll observe

  • Both features are documented and each works by itself
  • The failure appears at setup or on the first step rather than under memory pressure
  • Offloading was enabled precisely because memory was already tight, so simply removing it is not free
  • The error rarely names offloading, so the interaction is not obvious from the message

Common symptoms and what they mean

SymptomWhy it happens
The traceback ending inside HybridDeviceOptimizer while it builds its sub-optimizer parameter groups, not inside the sharding layerThe concrete mechanism in the Megatron implementation is worth stating exactly, because it explains why the error names the optimizer rather than the sharding. The offload path builds its host-side copy by calling detach, clone, cpu and pin_memory on each parameter in turn. Under sharding those parameters are no longer plain tensors but DTensors, so every one of those calls goes through a dispatcher that needs a sharding rule for the operator. Pinning host memory has no meaningful sharded semantics, the rule is absent, and the composition dies in the optimizer constructor.
A call to param.detach().clone().cpu().pin_memory() applied to a parameter that sharding has already replaced with a DTensorSharding divides a parameter across ranks and keeps authoritative placement metadata for each piece. Offloading relocates parameters to host memory and restores them on demand, keeping its own view of where each one is. Both are correct in isolation and each expects to be the authority, so composing them leaves two components disagreeing about a single tensor.
Frames passing through the DTensor dispatcher and its sharding propagator on the way to the failureThe disagreement shows up early because placement is resolved during setup and on the first use of a parameter, not when memory runs short. That is why the failure looks unrelated to the memory pressure that motivated enabling offloading.
The failure raised from get_megatron_optimizer during setup_model_and_optimizer, before the first training stepSupport for the combination is a property of a specific implementation and version rather than of the idea, so a stack that documents both features separately may not support them together, and a version that works is not evidence a neighbouring one does.
Setup or the first training step failing with both sharding and CPU offloading enabledThe concrete mechanism in the Megatron implementation is worth stating exactly, because it explains why the error names the optimizer rather than the sharding. The offload path builds its host-side copy by calling detach, clone, cpu and pin_memory on each parameter in turn. Under sharding those parameters are no longer plain tensors but DTensors, so every one of those calls goes through a dispatcher that needs a sharding rule for the operator. Pinning host memory has no meaningful sharded semantics, the rule is absent, and the composition dies in the optimizer constructor.
A device or placement error naming a parameter that offloading has moved to host memorySharding divides a parameter across ranks and keeps authoritative placement metadata for each piece. Offloading relocates parameters to host memory and restores them on demand, keeping its own view of where each one is. Both are correct in isolation and each expects to be the authority, so composing them leaves two components disagreeing about a single tensor.

Which systems are affected

  • Megatron-style FSDP implementations combined with CPU offloading
  • Stacks where offloading is configured by a launcher and sharding by the training framework, so neither knows about the other
  • Configurations carried over from a single-device memory-saving setup into a sharded one
  • Any setup where a parameter's owner is ambiguous between two memory managers

How to confirm this is the problem

Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.

  • ✓Read which frame raises. Landing in the optimizer's offload helper rather than in the sharding layer tells you offloading is the caller and sharding is only the thing it tripped over.
  • ✓Look for a DTensor dispatch frame beneath the offload call. Its presence means the parameter was already sharded when the host copy was attempted.
  • ✓Disable CPU offloading and re-run with sharding unchanged. Success isolates the interaction rather than either feature.
  • ✓Disable sharding and re-run with offloading unchanged. Success on both single-feature runs confirms the composition is the problem.
  • ✓Check whether the failure occurs at setup or on the first step rather than under sustained memory pressure, which distinguishes a placement disagreement from genuine exhaustion.

Root cause

  • The concrete mechanism in the Megatron implementation is worth stating exactly, because it explains why the error names the optimizer rather than the sharding. The offload path builds its host-side copy by calling detach, clone, cpu and pin_memory on each parameter in turn. Under sharding those parameters are no longer plain tensors but DTensors, so every one of those calls goes through a dispatcher that needs a sharding rule for the operator. Pinning host memory has no meaningful sharded semantics, the rule is absent, and the composition dies in the optimizer constructor.
  • Sharding divides a parameter across ranks and keeps authoritative placement metadata for each piece. Offloading relocates parameters to host memory and restores them on demand, keeping its own view of where each one is. Both are correct in isolation and each expects to be the authority, so composing them leaves two components disagreeing about a single tensor.
  • The disagreement shows up early because placement is resolved during setup and on the first use of a parameter, not when memory runs short. That is why the failure looks unrelated to the memory pressure that motivated enabling offloading.
  • Support for the combination is a property of a specific implementation and version rather than of the idea, so a stack that documents both features separately may not support them together, and a version that works is not evidence a neighbouring one does.

The fix and how to prevent it

Searchable error signature

search key
  File ".../megatron/core/optimizer/cpu_offloading/hybrid_optimizer.py", line 188, in _init_sub_optimizers
    ) = self._get_sub_optimizer_param_groups(self.offload_fraction)
    param = param.detach().clone().cpu().pin_memory()
  File ".../torch/distributed/tensor/_dispatch.py", line 160, in dispatch
    self.sharding_propagator.propagate(op_info)
Megatron FSDP does not work with CPU_OFFLOADING
RuntimeError raised during setup with both parameter sharding and CPU offloading enabled

Use this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.

The fix and the prevention pattern

The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.

Sign up free. Unlock the full analysis

No credit card. Daily allowance follows verified trust tier. Instant access.

Why the recommended fix works

Removing one claimant leaves a single authority and the contradiction disappears. Recovering the memory from activations works because activation memory is managed by a different mechanism entirely, so it does not compete for ownership of parameters. Widening sharding reduces the same footprint offloading was targeting, without a second manager.

Best practices by model family

Model / StackRecommendationNotes
Sharding must stayWiden sharding or shard optimizer state furtherReduces the same footprint without a second memory manager.
Memory pressure is from activationsActivation checkpointing or a smaller microbatchActivation memory does not compete for parameter ownership.
Offloading is essentialPin a version where the pair is supportedSupport is a property of the implementation and moves between releases.
DiagnosingTest each feature alone firstTwo successful single-feature runs prove it is the composition.

With the fix vs without the fix

DimensionWith the fixWithout the fix
What each feature controlsBoth control parameter placementAssumed to address different resources
When the failure appearsAt setup or the first stepExpected under sustained memory pressure
Whether documentation implies supportDocumented separately, not necessarily togetherRead as both being supported at once

Diagnostic note

“The instinct on hitting this is to look for a memory bug, because offloading was enabled under memory pressure and the failure arrives soon after. The timing is the clue that it is not: a genuine exhaustion appears when the workload grows, while a placement disagreement appears at setup or on the very first step regardless of size. Note which, before spending time on allocator settings.”

Visual fingerprint

Two managers, one parameter
  sharding    : parameter split across ranks, placement metadata per shard
  offloading  : parameter relocated to host memory, restored on demand

        both enabled ->  who owns the placement of this tensor?
                          |
                 resolved at setup / first use -> contradiction
Sharding and offloading each maintain their own record of where a parameter lives. With both enabled there is no agreed owner, and the contradiction surfaces the first time placement has to be resolved rather than when memory runs short.

Diagnose this failure in VS Code

Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.

Install the free VS Code extension

Frequently asked questions

Questions engineers and on-call staff commonly ask about this failure.

Both features are documented. Why can I not use both?
They are documented separately. Both control where a parameter lives, and support for composing them is a property of a specific implementation and version.
Is this an out-of-memory problem?
Usually not. A placement disagreement appears at setup or on the first step regardless of workload size, while genuine exhaustion appears as the workload grows.
I enabled offloading because I was out of memory.
Then recover it elsewhere: widen sharding, shard optimizer state further, or reduce activation memory, none of which claims ownership of parameters.
How do I confirm it is the combination?
Run each feature alone. Two successful single-feature runs and a failing combined run is the proof.

Don't just read the fix, diagnose your run

The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.