Tensor on device meta is not on the expected device: a device mismatch from a parameter never materialised
A model built with automatic device placement runs until one submodule executes with a parameter still on the placeholder device. Meta tensors carry shape and dtype but no data, so the operation refuses rather than computing on nothing, and the error names a device that is not a device at all.
A parameter was never given real data. Meta tensors are shape-only placeholders, and one of them reached a live operation. Find the submodule in the traceback, check it against the placement map, and place it explicitly.
- Symptom
RuntimeError: Tensor on device meta is not on the expected device cuda:0!- Root cause
- Loading a large model starts by building it without data. Every parameter is a meta tensor: correct shape, correct dtype, no storage. Placement then decides where each submodule will live and hooks are attached to bring the real weights in at the right moment.
- Recommended fix
- Identify the submodule in the traceback and check whether it appears in the placement map. Absence there is the entire fault and points at what to attach or place explicitly.
- How Denpex helps
- Denpex matches Tensor on device meta is not on the expected device: a device mismatch from a parameter never materialised across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
What this failure is
A forward-pass failure in which a parameter left on the meta placeholder device meets a real tensor in an operation, because automatic placement never attached a hook to materialise the submodule that owns it.
Is this what broke your run? Paste your log.
You're reading about Tensor on device meta is not on the expected device: a device mismatch from a parameter never materialised. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.
Why it happens (the mechanism)
Large-model loading deliberately separates structure from data so a model larger than one device can be described before it is populated. That separation is safe only while every parameter is subsequently materialised, and the mechanism responsible knows only about the modules it enumerated. Anything outside that enumeration keeps its placeholder, and a placeholder is indistinguishable from a real parameter until an operation needs its values.
What you'll observe
- Meta is not a real device, so the message reads as nonsense on first encounter
- Most of the model works, which points suspicion at the input rather than the weights
- The failing submodule is often a secondary encoder that placement treated differently from the main network
- The traceback ends in shape-inference internals, far from the placement decision that caused it
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| RuntimeError: Tensor on device meta is not on the expected device cuda:0 | Loading a large model starts by building it without data. Every parameter is a meta tensor: correct shape, correct dtype, no storage. Placement then decides where each submodule will live and hooks are attached to bring the real weights in at the right moment. A parameter that no hook covers stays a shape with nothing behind it. |
| The exception raised from elementwise shape-inference internals rather than from a kernel launch | The failure appears in shape inference rather than at a kernel because that layer runs first and it can see the contradiction immediately. One operand is on the accelerator, another exists nowhere, and there is no meaningful result to compute, so it raises instead of launching anything. |
| Forward passing through an offload hook wrapper immediately before the failing operation | The gap is usually structural rather than random. A submodule reached by a path placement did not walk, a component added to a pipeline after the map was built, or a parameter created during initialisation rather than loaded from the checkpoint will all be left behind by a mechanism that only knows about what it enumerated. |
| One submodule affected while the rest of the pipeline runs normally | Loading a large model starts by building it without data. Every parameter is a meta tensor: correct shape, correct dtype, no storage. Placement then decides where each submodule will live and hooks are attached to bring the real weights in at the right moment. A parameter that no hook covers stays a shape with nothing behind it. |
Which systems are affected
- Automatic device placement, which loads a skeleton on the placeholder device and fills it in per submodule
- Offload hooks that move weights onto the accelerator immediately before a submodule runs and off it afterwards
- Multi-component pipelines where a vision or text encoder is placed separately from the main network
- Modules constructed after placement was computed, which no hook was ever attached to
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Walk the model's parameters and buffers after loading and report any whose device is the placeholder. The list names exactly what was missed.
- ✓Compare the failing submodule's name against the placement map. A submodule with no entry was never scheduled to be materialised.
- ✓Load the same model onto a single device without automatic placement. Running correctly there confirms placement rather than the model or the input.
Root cause
- Loading a large model starts by building it without data. Every parameter is a meta tensor: correct shape, correct dtype, no storage. Placement then decides where each submodule will live and hooks are attached to bring the real weights in at the right moment. A parameter that no hook covers stays a shape with nothing behind it.
- The failure appears in shape inference rather than at a kernel because that layer runs first and it can see the contradiction immediately. One operand is on the accelerator, another exists nowhere, and there is no meaningful result to compute, so it raises instead of launching anything.
- The gap is usually structural rather than random. A submodule reached by a path placement did not walk, a component added to a pipeline after the map was built, or a parameter created during initialisation rather than loaded from the checkpoint will all be left behind by a mechanism that only knows about what it enumerated.
The fix and how to prevent it
Searchable error signature
RuntimeError: Tensor on device meta is not on the expected device cuda:0!
File ".../torch/_prims/__init__.py", in _prim_elementwise_meta
File ".../torch/_library/fake_impl.py", in meta_kernel
File ".../accelerate/hooks.py", in new_forwardUse this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Why the recommended fix works
Placing the component explicitly gives its parameters real storage on a real device, so the operation has values on both sides and proceeds. Materialising the whole model achieves the same by leaving no placeholders anywhere. Neither changes how the model computes; they change whether the numbers exist at the moment they are required.
Best practices by model family
| Model / Stack | Recommendation | Notes |
|---|---|---|
| Single component missed | Place that component explicitly | Cleanest when the rest of the pipeline is correct. |
| Model fits on one device | Materialise fully at load | No placeholders survive, so none can be reached. |
| Sequential pipeline offload | Register every component | Later additions are frequently not registered. |
| Parameters made at init | Materialise them after construction | A checkpoint-derived map cannot see them. |
With the fix vs without the fix
| Dimension | With the fix | Without the fix |
|---|---|---|
| What meta means | Shape and dtype with no storage | Read as an unusual accelerator |
| Which operand is at fault | The weight, always | Suspected to be the input |
| When it is detectable | Immediately after load, by traversal | Only when the module runs |
Diagnostic note
“The instinct is to look at the inputs, because the error mentions a device and inputs are what usually arrive on the wrong one. It is worth resisting for one minute and reading which side of the operation is on meta: an input can be misplaced, but it is never on meta, because inputs are made of data. Meta on either operand means a weight, and a weight on meta means placement, every time.”
Visual fingerprint
build skeleton every parameter on meta (shape only)
|
v
compute placement map enumerates modules it can see
|
v
attach hooks -> materialise on use
|
+--> covered module : real weights ✓
+--> module not in map : still on meta ✗ -> RuntimeError on forwardDiagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionFrequently asked questions
Questions engineers and on-call staff commonly ask about this failure.
What is device meta?
Is my input on the wrong device?
Why does most of the model work?
How do I catch this earlier?
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.