Skip to content

Operator task guide

Fix the device mismatch at the failing PyTorch operation

The same-device error means an operation received incompatible device placement. It does not mean every tensor in the program belongs on cuda:0. Inspect the operands at the failing line and choose the device owned by that model or local rank.

Check newly created tensors, registered buffers and restored optimizer state as well as model inputs. Tensor conversion returns a tensor; an ignored return value can leave the original on CPU. Indexing rules are operation-specific.

Reviewed . Reference guidance is not a diagnosis of your workload.

Step-by-step method

  1. 1

    Inspect only the failing operands

    At the first traceback line in your code, record each participating tensor device, dtype and shape. Do not print private tensor contents. For distributed runs include local rank and the visible-device mapping.

    # Replace x and y with the operands at your failing line.
    for name, value in [('x', x), ('y', y)]:
        print(name, value.device, value.dtype, tuple(value.shape))
  2. 2

    Check the owning model and tensor creation

    Compare inputs with the relevant parameter and registered buffer devices. Use factories such as zeros_like when inheriting placement is intended, or assign the returned tensor conversion to a variable. Do not move a sharded model indiscriminately to one GPU.

  3. 3

    Check indices and resumed state separately

    Inspect the operation documentation for index dtype and placement, rather than applying one rule to all indexing. If forward and backward work but optimizer.step fails after resume, inspect the relevant optimizer state tensors and use the framework-supported restore flow.

  4. 4

    Prove the corrected operation is still correct

    Rerun the smallest affected batch. Compare output shape and values with a known reference. For training also check loss, intended gradients and parameter updates. Then exercise the resumed or distributed path that originally failed, not only a new single-device run.

Where placement can diverge

Signals, meanings and actions for fix the device mismatch at the failing pytorch operation.
SignalWhat it meansNext action
CPU input with CUDA parameterThe forward operation has incompatible operands.Transfer the input to the owning model device and retain the returned tensor.
Tensor created inside forwardA default factory may allocate on CPU.Choose placement and dtype from the relevant input or parameter.
Failure at indexing or gatherIndex device and dtype rules depend on that operation.Inspect indices at that line and consult the exact operator contract.
Only resumed optimizer.step failsRestored state may not match parameter placement.Inspect state for that optimizer and follow its supported restore sequence.

Evidence checklist

  • First traceback and failing expression
  • Operand device, dtype and shape
  • Relevant parameter and buffer placement
  • Checkpoint restore order if applicable
  • Local rank and visible-device mapping if applicable

Common mistakes

Calling cuda() on everything

This can break CPU-owned indices, model sharding or the local-rank device assignment.

Checking only for an absent exception

A different input or a new optimizer can hide the original defect. Verify the original workload path and intended values.

Frequently asked questions

What does indices should be either on cpu or on the same device as the indexed tensor mean?

Inspect the index tensor and the tensor being indexed at that exact operation. This message permits CPU indices or indices on the indexed tensor device. It does not say all index operators have that contract, or that every tensor should be moved to cuda:0.

Should every index tensor be moved to CUDA?

No universal rule applies. Advanced indexing and individual index operators have different contracts. Use the failing operation and its documentation to choose dtype and placement.

Why does tensor.to(device) appear not to work?

A conversion can return a different tensor. Retain that return value. Also confirm you are checking the tensor actually passed to the failing operation.

Apply the method to your incident

Use the three free diagnoses to review your error and relevant evidence. Keep reference guidance separate from the cause and recovery status of your own workload.

Diagnose your incident