Calling cuda() on everything
This can break CPU-owned indices, model sharding or the local-rank device assignment.
Operator task guide
The same-device error means an operation received incompatible device placement. It does not mean every tensor in the program belongs on cuda:0. Inspect the operands at the failing line and choose the device owned by that model or local rank.
Check newly created tensors, registered buffers and restored optimizer state as well as model inputs. Tensor conversion returns a tensor; an ignored return value can leave the original on CPU. Indexing rules are operation-specific.
Reviewed . Reference guidance is not a diagnosis of your workload.
At the first traceback line in your code, record each participating tensor device, dtype and shape. Do not print private tensor contents. For distributed runs include local rank and the visible-device mapping.
# Replace x and y with the operands at your failing line.
for name, value in [('x', x), ('y', y)]:
print(name, value.device, value.dtype, tuple(value.shape))Compare inputs with the relevant parameter and registered buffer devices. Use factories such as zeros_like when inheriting placement is intended, or assign the returned tensor conversion to a variable. Do not move a sharded model indiscriminately to one GPU.
Inspect the operation documentation for index dtype and placement, rather than applying one rule to all indexing. If forward and backward work but optimizer.step fails after resume, inspect the relevant optimizer state tensors and use the framework-supported restore flow.
Rerun the smallest affected batch. Compare output shape and values with a known reference. For training also check loss, intended gradients and parameter updates. Then exercise the resumed or distributed path that originally failed, not only a new single-device run.
| Signal | What it means | Next action |
|---|---|---|
| CPU input with CUDA parameter | The forward operation has incompatible operands. | Transfer the input to the owning model device and retain the returned tensor. |
| Tensor created inside forward | A default factory may allocate on CPU. | Choose placement and dtype from the relevant input or parameter. |
| Failure at indexing or gather | Index device and dtype rules depend on that operation. | Inspect indices at that line and consult the exact operator contract. |
| Only resumed optimizer.step fails | Restored state may not match parameter placement. | Inspect state for that optimizer and follow its supported restore sequence. |
This can break CPU-owned indices, model sharding or the local-rank device assignment.
A different input or a new optimizer can hide the original defect. Verify the original workload path and intended values.
Inspect the index tensor and the tensor being indexed at that exact operation. This message permits CPU indices or indices on the indexed tensor device. It does not say all index operators have that contract, or that every tensor should be moved to cuda:0.
No universal rule applies. Advanced indexing and individual index operators have different contracts. Use the failing operation and its documentation to choose dtype and placement.
A conversion can return a different tensor. Retain that return value. Also confirm you are checking the tensor actually passed to the failing operation.
Use the three free diagnoses to review your error and relevant evidence. Keep reference guidance separate from the cause and recovery status of your own workload.
Diagnose your incident