Multi-Allocator Zero-Copy MMU Fault
A secondary memory allocator (like CuPy's memory pool) forcibly freed a block of memory that was zero-copied from PyTorch. PyTorch attempts to access the now-unmapped virtual address.
A secondary memory allocator (like CuPy's memory pool) forcibly freed a block of memory that was zero-copied from PyTorch.
- Root cause
- A secondary memory allocator (like CuPy's memory pool) forcibly freed a block of memory that was zero-copied from PyTorch. PyTorch attempts to access the now-unmapped virtual address.
- Recommended fix
- Enforce explicit tensor cloning across allocator boundaries. - tensor_cupy = cp.asarray(tensor_torch.clone().detach())
- How Denpex helps
- Denpex matches Multi-Allocator Zero-Copy MMU Fault across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
What this failure is
Multi-Allocator Zero-Copy MMU Fault is a CUDA failure seen during ML training runs. A secondary memory allocator (like CuPy's memory pool) forcibly freed a block of memory that was zero-copied from PyTorch. PyTorch attempts to access the now-unmapped virtual address. Common tags: Memory Allocator, User Report.
Is this what broke your run? Paste your log.
You're reading about Multi-Allocator Zero-Copy MMU Fault. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.
Want 14 days on the Scale plan?
Request an evaluation code. A verified workplace organization activates up to 50 diagnoses a day, alerts, history, and follow-up questions. No credit card or automatic subscription.
Why it happens (the mechanism)
An Xid 31 MMU fault is often blamed on physical hardware defects, bad VRAM, or driver bugs. In reality, it is a strict software-level use-after-free vulnerability across library boundaries.
What you'll observe
- The entire CUDA context dies instantly.
- The training pipeline resets or crashes completely.
- Memory dumps show the faulting address in high-aligned memory spaces (e.g., 0x725f_fb800000).
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| NVRM: Xid (PCI:0000:XX:00): 31, pid=XXXX, name=python, MMU Fault: ENGINE GRAPHICS | A secondary memory allocator (like CuPy's memory pool) forcibly freed a block of memory that was zero-copied from PyTorch. PyTorch attempts to access the now-unmapped virtual address. |
| Fault is of type FAULT_PDE | A secondary memory allocator (like CuPy's memory pool) forcibly freed a block of memory that was zero-copied from PyTorch. PyTorch attempts to access the now-unmapped virtual address. |
Which systems are affected
- PyTorch
- CuPy
- TensorRT
- CUDA
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Decode the Xid 31 payload to identify if it is a FAULT_PDE (wholesale deallocation) or FAULT_PTE (single page issue).
- ✓Trace the tensor lifecycle using cuda-memcheck or Compute Sanitizer.
- ✓Identify zero-copy handoffs (e.g., DLPack operations) in the code.
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Diagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionRoot cause
- A secondary memory allocator (like CuPy's memory pool) forcibly freed a block of memory that was zero-copied from PyTorch. PyTorch attempts to access the now-unmapped virtual address.
The fix and how to prevent it
Evaluate Denpex on your own logs
Request a Scale evaluation code. A verified workplace organization activates 14 days with up to 50 diagnoses a day. Every account keeps its current diagnosis allowance and gets a verification path. No card or automatic subscription.
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.