OFI Memlock Exhaustion. RDMA Memory Registration Failure
NCCL's OFI (OpenFabrics Interface) transport fails to register memory for RDMA when the OS memlock limit is too low. Training crashes at NCCL init with 'Unable to register memory RC:12'. Denpex detects the insufficient memlock limit from NCCL logs and recommends the system-level fix.
NCCL's OFI (OpenFabrics Interface) transport fails to register memory for RDMA when the OS memlock limit is too low.
What this failure is
OFI Memlock Exhaustion. RDMA Memory Registration Failure is a Communication failure seen during ML training runs. NCCL's OFI (OpenFabrics Interface) transport fails to register memory for RDMA when the OS memlock limit is too low. Training crashes at NCCL init with 'Unable to register memory RC:12'. Denpex detects the insufficient memlock limit from NCCL logs and recommends the system-level fix. Common tags: Ofi, Memlock, Rdma, Memory Registration.
Is this what broke your run? Paste your log.
You're reading about OFI Memlock Exhaustion. RDMA Memory Registration Failure. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.
Want 14 days on the Scale plan?
Request an evaluation code. A verified workplace organization activates up to 50 diagnoses a day, alerts, history, and follow-up questions. No credit card or automatic subscription.
Why it happens (the mechanism)
The OS per-process memlock limit (ulimit -l) is too low for RDMA memory registration. RDMA requires pinning (locking) GPU memory in physical RAM so the NIC can DMA directly to it. The default memlock limit on many Linux distributions is 64KB, which is far too low for NCCL's RDMA buffers. Containers inherit the host's memlock limit unless explicitly overridden. Taken together, these mechanisms explain why the failure is reproducible, why it tends to surface on specific workloads or scales, and why generic mitigation attempts often fall short without addressing the underlying cause.
What you'll observe
- NCCL initialization fails with 'Unable to register memory' on the OFI transport
- Training works with NCCL_NET=Socket but fails with the default OFI/IB transport
- The error appears on all ranks simultaneously during NCCL channel setup
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| NCCL WARN NET/OFI Unable to register memory RC:12 | The OS per-process memlock limit (ulimit -l) is too low for RDMA memory registration |
| NCCL init timeout after the memory registration error | RDMA requires pinning (locking) GPU memory in physical RAM so the NIC can DMA directly to it |
| Training works when switching to Socket transport: NCCL_NET=Socket | The default memlock limit on many Linux distributions is 64KB, which is far too low for NCCL's RDMA buffers |
| ulimit -l shows a low memlock value (e.g., 64KB or 65536) | Containers inherit the host's memlock limit unless explicitly overridden |
Which systems are affected
- Multi-node training with AWS EFA (OFI transport)
- Any RDMA-based NCCL transport on Linux
- Containers without proper memlock configuration
- Fresh OS installations where memlock defaults to 64KB
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Reproduce the failure from a clean checkpoint/seed: the symptom must appear without warm-up state from a previous run.
- ✓Verified signal present: NCCL WARN NET/OFI Unable to register memory RC:12
- ✓Verified signal present: NCCL init timeout after the memory registration error
- ✓Verified signal present: Training works when switching to Socket transport: NCCL_NET=Socket
- ✓Verified signal present: ulimit -l shows a low memlock value (e.g., 64KB or 65536)
- ✓A targeted fix from the "How to fix it" section eliminates or substantially reduces the symptom within one validation pass.
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Diagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionNCCL errors in context
NCCL is where a distributed job reports failure, which is not the same as where it failed. The hub lists every common NCCL error next to what it actually indicates, and the environment variables that tell them apart.
Compare every nccl error side by sideRelated failures to investigate next
Root cause
- The OS per-process memlock limit (ulimit -l) is too low for RDMA memory registration
- RDMA requires pinning (locking) GPU memory in physical RAM so the NIC can DMA directly to it
- The default memlock limit on many Linux distributions is 64KB, which is far too low for NCCL's RDMA buffers
- Containers inherit the host's memlock limit unless explicitly overridden
The fix and how to prevent it
Evaluate Denpex on your own logs
Request a Scale evaluation code. A verified workplace organization activates 14 days with up to 50 diagnoses a day. Every account keeps its current diagnosis allowance and gets a verification path. No card or automatic subscription.
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.