Skip to content

Operator task guide

Use Compute Sanitizer on the smallest failing CUDA workload

Compute Sanitizer memcheck can attribute out-of-bounds and misaligned GPU memory accesses. Start with a small reproduction of the first failing operation, not the entire production training run.

Instrumentation can change timing and add substantial overhead. A finding is evidence about that reproduction. A clean run does not prove that every input, race or hardware path is healthy.

Reviewed . Reference guidance is not a diagnosis of your workload.

Step-by-step method

  1. 1

    Reduce the case without removing the failure

    Preserve the offending shapes, dtype, kernel selection and inputs using a safe synthetic equivalent where possible. Confirm the unsanitized reproduction still fails. Work outside the production service.

  2. 2

    Run the matching tool and retain its exit status

    Use memcheck for memory access faults. Use an installed CUDA toolkit version compatible with the workload. The example assumes your minimal program is repro.py; it is not a flag to add to an unrelated training launcher.

    compute-sanitizer --tool memcheck --error-exitcode 99 python repro.py
  3. 3

    Attribute the first finding

    Save the first invalid access and its kernel/source context, not just the final runtime exception. Compile your own CUDA extension with line information when applicable. If worker subprocesses are involved, consult the tool documentation for the target process scope.

  4. 4

    Verify both the fix and intended values

    Rerun the reproduction with the same input, first under the tool and then without instrumentation. Compare outputs or gradients to an independent reference, then test representative production shapes. Keep unresolved findings open.

Tool findings and their limits

Signals, meanings and actions for use compute sanitizer on the smallest failing cuda workload.
SignalWhat it meansNext action
Invalid global read or writeA memory access in the observed kernel is invalid.Check index bounds, allocation lifetime and the source location.
Synchronization or shared-memory concernMemory checking alone may not test the relevant hazard.Choose synccheck or racecheck for the matching documented scope.
No findings on a reduced caseNo covered error was detected for that execution.Check that the reduction retained the failure and test other affected inputs.

Evidence checklist

  • Exact reproduction and tool command
  • CUDA/toolkit, driver and extension versions
  • Shapes, dtype and kernel selection
  • First sanitizer finding with attribution
  • Output or gradient comparison after correction

Common mistakes

Running it against a shared live service

Instrumentation can heavily slow work and change timing. Use an isolated reproduction.

Calling a clean tool result a hardware certification

The check covers particular code paths and inputs. It does not certify the GPU or all workload behavior.

Frequently asked questions

Does memcheck prove the training result is correct?

No. It checks covered memory accesses. Verify intended outputs, gradients and updates separately.

What does error-exitcode 99 do?

It makes a detected tool error produce the selected nonzero exit status when the application itself succeeded, so a check does not silently count that run as passing.

Apply the method to your incident

Use the three free diagnoses to review your error and relevant evidence. Keep reference guidance separate from the cause and recovery status of your own workload.

Diagnose your incident