FSDP preparation fails because activation checkpointing calls a wrapping policy that was never set
Choosing not to auto-wrap is legitimate and leaves the wrapping policy unset. Activation checkpointing then tries to apply that policy to decide which submodules to checkpoint, calls the empty value, and preparation fails with a type error naming nothing that appears in the configuration.
The two settings are not independent. Activation checkpointing calls the wrapping policy, and no-wrap leaves it empty. Set an explicit policy with a layer class, or turn activation checkpointing off.
- Symptom
TypeError: 'NoneType' object is not callable- Root cause
- Selecting no automatic wrapping sets the policy to an empty value, which is correct: there is no policy because nothing is to be wrapped automatically. Activation checkpointing, however, reuses that policy to decide which submodules receive checkpointing, and it assumes something callable is there. The two settings are presented as independent and are not.
- Recommended fix
- Choose an explicit wrapping policy and name the layer class to wrap, which gives activation checkpointing something to apply and is the configuration these frameworks are designed around.
- How Denpex helps
- Denpex matches FSDP preparation fails because activation checkpointing calls a wrapping policy that was never set across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
What this failure is
A preparation-time type error in which activation checkpointing invokes an FSDP wrapping policy that is legitimately unset, because the configuration selected manual wrapping while leaving checkpointing enabled.
Is this what broke your run? Paste your log.
You're reading about FSDP preparation fails because activation checkpointing calls a wrapping policy that was never set. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.
Why it happens (the mechanism)
Configuration formats present orthogonal-looking switches, and the code behind them shares state. Selecting manual wrapping correctly produces no policy; activation checkpointing then reuses that policy as its own selector of what to checkpoint, without checking that one exists. The result is a permitted configuration that no code path can serve.
What you'll observe
- The error names a type rather than a setting, so it does not point at the configuration that caused it
- Each of the two options works on its own and only the combination fails
- The traceback lands in framework internals with no application frame to act on
- Turning off the wrong one of the two makes it pass and teaches the wrong lesson
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| TypeError: 'NoneType' object is not callable raised during model preparation | Selecting no automatic wrapping sets the policy to an empty value, which is correct: there is no policy because nothing is to be wrapped automatically. Activation checkpointing, however, reuses that policy to decide which submodules receive checkpointing, and it assumes something callable is there. |
| Traceback through accelerator.prepare into the FSDP2 preparation path and an activation-checkpointing helper | The two settings are presented as independent and are not. One expresses an intent about sharding granularity and the other silently consumes it, so a combination the configuration format permits is one the code path cannot serve. |
| A configuration selecting no automatic wrapping together with activation checkpointing enabled | A neighbouring failure produces the same empty policy from the other direction: a transformer-based policy on a model that exposes no splittable module list, or with a class name that cannot be resolved, also yields nothing to call. |
| Could not find the transformer layer class to wrap in the model, when a class name was supplied instead | Selecting no automatic wrapping sets the policy to an empty value, which is correct: there is no policy because nothing is to be wrapped automatically. Activation checkpointing, however, reuses that policy to decide which submodules receive checkpointing, and it assumes something callable is there. |
Which systems are affected
- Accelerate and similar launchers where wrapping policy and activation checkpointing are independent switches
- Custom models outside a standard architecture library, which expose no default list of splittable modules
- Configurations converted from an older setup where activation checkpointing was enabled for other reasons
- Parameter-efficient tuning stacks that assemble the wrapping policy themselves
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Check whether activation checkpointing is enabled while the wrapping policy is set to no automatic wrapping; that pairing is the failure.
- ✓Disable activation checkpointing alone and re-run. Success identifies it as the consumer of the absent policy.
- ✓Where a transformer-based policy is configured, confirm the named layer class resolves against the model, since an unresolvable name leaves the policy empty too.
Root cause
- Selecting no automatic wrapping sets the policy to an empty value, which is correct: there is no policy because nothing is to be wrapped automatically. Activation checkpointing, however, reuses that policy to decide which submodules receive checkpointing, and it assumes something callable is there.
- The two settings are presented as independent and are not. One expresses an intent about sharding granularity and the other silently consumes it, so a combination the configuration format permits is one the code path cannot serve.
- A neighbouring failure produces the same empty policy from the other direction: a transformer-based policy on a model that exposes no splittable module list, or with a class name that cannot be resolved, also yields nothing to call.
The fix and how to prevent it
Searchable error signature
TypeError: 'NoneType' object is not callable
File ".../accelerate/accelerator.py", in _prepare_fsdp2
File ".../accelerate/utils/fsdp_utils.py", in fsdp2_apply_ac
ValueError: Could not find the transformer layer class to wrap in the model.Use this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Why the recommended fix works
An explicit policy gives both consumers something real: sharding gets its granularity and checkpointing gets its selector. Disabling checkpointing works by removing the consumer instead. Either restores the invariant the code assumes, which is that a policy exists whenever something wants to apply one.
Best practices by model family
| Model / Stack | Recommendation | Notes |
|---|---|---|
| Standard transformer architecture | Transformer-based policy naming the decoder block | Gives sharding and checkpointing the same real unit. |
| Custom model with no module list | Name the layer class explicitly | Defaults derived from the model are absent, so the policy stays empty. |
| No natural layer class | Size-based policy with a minimum parameter count | Produces a real policy without naming an architecture. |
| Wrapping must stay manual | Disable activation checkpointing | It is the component consuming the policy that does not exist. |
With the fix vs without the fix
| Dimension | With the fix | Without the fix |
|---|---|---|
| Are the two settings independent | No, checkpointing consumes the wrapping policy | Presented as separate switches |
| What an empty policy means | A valid expression of manual wrapping | Read as a misconfiguration in itself |
| Which switch to change | The wrapping policy | Whichever one makes the error stop |
Diagnostic note
“The trap is that disabling either option makes the error disappear, so whichever one you try first looks like the culprit. Disabling activation checkpointing is usually the wrong lesson to take away, because it silently costs memory on every subsequent run. Set the wrapping policy properly; it is the setting that was actually incomplete.”
Visual fingerprint
auto_wrap_policy = NO_WRAP -> policy = None (legitimate)
|
+------------------------+------------------------+
v v
sharding: nothing to wrap, fine activation checkpointing: policy(...)
-> None is not callableDiagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionRelated failures to investigate next
Frequently asked questions
Questions engineers and on-call staff commonly ask about this failure.
Both settings are valid on their own. Why does the pair fail?
Should I just turn off activation checkpointing?
I set a transformer policy and still get an empty policy.
Which layer class should I name?
References
- ↗Accelerate issue #3822: no-wrap policy combined with activation checkpointing raising a type error
- ↗Accelerate issue #2947: how layers are checked when wrapping for FSDP
- ↗PEFT issue #2166: the wrapping policy failing when no splittable module list exists
- ↗Transformers documentation on FSDP wrapping policies and layer classes
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.