Skip to content

Multi-Allocator Zero-Copy MMU Fault

A secondary memory allocator (like CuPy's memory pool) forcibly freed a block of memory that was zero-copied from PyTorch. PyTorch attempts to access the now-unmapped virtual address.

Quick answer

A secondary memory allocator (like CuPy's memory pool) forcibly freed a block of memory that was zero-copied from PyTorch.

Root cause
A secondary memory allocator (like CuPy's memory pool) forcibly freed a block of memory that was zero-copied from PyTorch. PyTorch attempts to access the now-unmapped virtual address.
Recommended fix
Enforce explicit tensor cloning across allocator boundaries. - tensor_cupy = cp.asarray(tensor_torch.clone().detach())
How Denpex helps
Denpex matches Multi-Allocator Zero-Copy MMU Fault across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
CUDA#memory allocator#user-report

What this failure is

Multi-Allocator Zero-Copy MMU Fault is a CUDA failure seen during ML training runs. A secondary memory allocator (like CuPy's memory pool) forcibly freed a block of memory that was zero-copied from PyTorch. PyTorch attempts to access the now-unmapped virtual address. Common tags: Memory Allocator, User Report.

Live diagnosis, no signup

Is this what broke your run? Paste your log.

You're reading about Multi-Allocator Zero-Copy MMU Fault. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.

training_logs.txt
No log to hand? Try one:

3 free diagnoses/day

Want 14 days on the Scale plan?

Request an evaluation code. A verified workplace organization activates up to 50 diagnoses a day, alerts, history, and follow-up questions. No credit card or automatic subscription.

Evaluate one incident

Why it happens (the mechanism)

An Xid 31 MMU fault is often blamed on physical hardware defects, bad VRAM, or driver bugs. In reality, it is a strict software-level use-after-free vulnerability across library boundaries.

What you'll observe

  • The entire CUDA context dies instantly.
  • The training pipeline resets or crashes completely.
  • Memory dumps show the faulting address in high-aligned memory spaces (e.g., 0x725f_fb800000).

Common symptoms and what they mean

SymptomWhy it happens
NVRM: Xid (PCI:0000:XX:00): 31, pid=XXXX, name=python, MMU Fault: ENGINE GRAPHICSA secondary memory allocator (like CuPy's memory pool) forcibly freed a block of memory that was zero-copied from PyTorch. PyTorch attempts to access the now-unmapped virtual address.
Fault is of type FAULT_PDEA secondary memory allocator (like CuPy's memory pool) forcibly freed a block of memory that was zero-copied from PyTorch. PyTorch attempts to access the now-unmapped virtual address.

Which systems are affected

  • PyTorch
  • CuPy
  • TensorRT
  • CUDA

How to confirm this is the problem

Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.

  • Decode the Xid 31 payload to identify if it is a FAULT_PDE (wholesale deallocation) or FAULT_PTE (single page issue).
  • Trace the tensor lifecycle using cuda-memcheck or Compute Sanitizer.
  • Identify zero-copy handoffs (e.g., DLPack operations) in the code.

The fix and the prevention pattern

The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.

Sign up free. Unlock the full analysis

No credit card. Daily allowance follows verified trust tier. Instant access.

Diagnose this failure in VS Code

Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.

Install the free VS Code extension

Root cause

  • A secondary memory allocator (like CuPy's memory pool) forcibly freed a block of memory that was zero-copied from PyTorch. PyTorch attempts to access the now-unmapped virtual address.

The fix and how to prevent it

Evaluate Denpex on your own logs

Request a Scale evaluation code. A verified workplace organization activates 14 days with up to 50 diagnoses a day. Every account keeps its current diagnosis allowance and gets a verification path. No card or automatic subscription.

We send a single-use code tied to that address. Static provider and TLD rules do not reject valid addresses. Account trust determines the benefit after signup.

Don't just read the fix, diagnose your run

The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.