A 32-node DDP job ends in an NCCL timeout after one rank exhausts memory. The useful path is to identify the first failed rank, prove the memory growth, and change the responsible configuration before investigating the fabric.
When one rank dies, all you see is 512 NCCL timeouts.
Denpex separates the initiating event from the cascade, identifies the rank when the evidence supports it, classifies the failure domain, and gives you a concrete next action. Built for distributed PyTorch, FSDP, and DeepSpeed at 64 to 16,384 GPUs.
Paste a crash log for a real diagnosis. No signup, no install. Or explore the live console →
Animated demo. A 512-GPU llama3-70b fine-tuning run crashes. Denpex reads the logs and reports: symptom, 512 NCCL collective timeouts that were only collateral; cause, rank 42 on node-05, GPU 8A:00, Xid 48 double-bit ECC error; class, hardware, route to infra rather than ML. It recommends draining node-05 and resuming with elastic restarts from step 84,000 with no work lost, then alerts the on-call engineer with the fix attached.
Built for
Frontier LLM training
70B+ model training across 1,000+ GPUs. Cross-rank cascade analysis, silent data corruption detection, and resume from the last good checkpoint.
FSDP / DDP production
PyTorch DDP and FSDP at scale. Diagnose the originating rank, the failure class, and the evidence-backed next action.
GPU clouds and neoclouds
White-label diagnosis for your customers. Offer Denpex under your brand, on the clusters you already operate.
Healthcare and life sciences
HIPAA BAA on Data Center. Configurable PHI masking on the agent, with audit logs of every diagnosis.
Evaluate Denpex with your incidents. Scope a pilot →
Built for distributed ML training.
From frontier LLM training to FSDP fine-tunes to GPU cloud operations. Denpex fits the way your team already runs.
Frontier LLM training
70B+ model training across 1,000+ GPUs. Cross-rank cascade analysis, silent data corruption detection, and resume from the last good checkpoint.
FSDP / DDP production
PyTorch DDP and FSDP at scale. Diagnose the originating rank, the failure class, and the evidence-backed next action.
GPU clouds and neoclouds
White-label diagnosis for your customers. Offer Denpex under your brand, on the clusters you already operate.
Healthcare and life sciences
HIPAA BAA on Data Center. Configurable PHI masking on the agent, with audit logs of every diagnosis.
Supported Ecosystems
Frameworks
Schedulers
Monitoring
Who uses Denpex?
Primary Operators
Training Infrastructure Engineers, ML Infrastructure Engineers, AI Platform Engineers, MLOps Engineers, Distributed Systems Engineers, GPU Cluster Engineers, SREs, and Platform Reliability Engineers.
Secondary Beneficiaries
AI Researchers, Machine Learning Engineers, Deep Learning Engineers, Data Scientists, and Performance Engineers.
Platform Buyers
Director of AI Infrastructure, Head of ML Platform, VP of Engineering, CTO, and Director of Research Infrastructure.
Ideal Organizations
AI startups training foundation models, Enterprise AI teams, GPU cloud providers, Managed ML platforms, Universities, National laboratories (HPC), Autonomous vehicle companies, and Defense contractors.
Your logs say “NCCL timeout.” The timeline says rank 17.
Distributed timeline reconstruction orders every event across every rank, clock-drift corrected, so the cascade reads in causal order, not log order. The watchdog is the last thing that happened, never the first.
The complete reliability surface
Diagnosis is the entry point. The platform covers the whole failure lifecycle before launch, during training, and after the fix ships.
10,800+ deterministic signatures plus retrieval over 700+ published Failure Encyclopedia entries, with deeper fallback analysis for anything novel. Every diagnosis ends in a concrete next action.
Root cause analysis
The best-supported initiating fault, separated from the symptom that woke you up
First-failed-rank detection
The initiating rank when the submitted evidence supports attribution; unknown otherwise
Distributed rank correlation
Cross-rank telemetry stitched into one causal picture
Cascade failure analysis
How one bad GPU took 63 healthy ranks down with it
NCCL timeout diagnosis
The initiator behind the watchdog's generic timeout
CUDA OOM diagnosis + memory attribution
The tensor or layer when allocator or profiler evidence identifies it; otherwise the next discriminator
Memory fragmentation diagnosis
Reserved-but-unallocated signatures, allocator-level fixes
Gradient explosion diagnosis
Norm spikes traced back to layer and step
NaN loss diagnosis
The propagation source, not just the first poisoned batch
Weight divergence diagnosis
Drift measured against your own healthy baselines
Silent hang diagnosis
Heartbeat detection for jobs that die without a stack trace
Device assert diagnosis
Device-side asserts mapped to the offending operation
Checkpoint corruption diagnosis
Torn writes and truncated shards caught before resume
Import error diagnosis
Environment faults separated from training faults
Version mismatch diagnosis
PyTorch x CUDA x cuDNN conflicts flagged precisely
Disk full diagnosis
Storage exhaustion before it masquerades as a framework crash
AI fallback analysis
Unknown failures get deep analysis on masked excerpts
Prescriptive fixes
Copy-paste resolution paths, verified against the failure class
Resume checkpoint recommendations
The last verified-good step to restart from
Hardware vs software classification
Infra issue or ML issue, instantly, so the right team moves
Looking for the full encyclopedia? Browse 700+ published entries →
Architecture your security team can say yes to.
No paste-your-logs surprises. The agent is transparent about what it touches, what it masks, and what (if anything) leaves your cluster.
In-VPC agent
Single Python file, stdlib only, no root, no kernel module. Wraps your training command, heartbeats every 120 s. PII/PHI masking runs client-side, before any byte leaves your cluster, on by default. Set DENPEX_PRIVACY=strict and the agent ships only anonymized failure signatures; raw logs never leave.
Deterministic engine
10,800+ deterministic regex signatures plus IDF-weighted retrieval over 700+ published Failure Encyclopedia entries, with clock-drift-corrected timeline reconstruction, do the work deterministically. The AI fallback only sees masked excerpts of novel failures, and your logs are never used to train anything.
Routed resolution
Ownership mapping sends one correlated incident, evidence-ranked cause, classification, and concrete next action to the engineer who owns the job, on Slack, PagerDuty, SMS or webhook. Hardware issues route to infra; ML issues route to research.
What your failures cost. And what you get back.
Estimate the GPU-hours and engineering time your team loses to undiagnosed failures, and the payback on each plan.
The kind of failure this catches
Illustrative incident paths based on documented failure modes in the Denpex encyclopedia. They are not customer quotes or claims about a specific engagement.
An FSDP fine-tune fails only when a particular sequence shape reaches a corrupted sample. Reproducing the same sample on a known-good node separates the data path from an unnecessary hardware escalation.
A distributed job reports a collective timeout after Rank 47 hits a CUDA OOM. A shared evidence trail lets the model and infrastructure teams work from the initiating failure instead of treating the final timeout as the cause.
Priced against your GPU bill, not your seat count.
A single failure on a 64-GPU cluster wastes hours of compute and an afternoon of engineering time. Every plan starts free, no credit card needed. Annual plans save 2 months.
Not sure which plan? Match your monthly GPU-hours.
- 0, 25,000 GPU-hrs/moTeam
- 25,000, 150,000 GPU-hrs/moScale
- 150,000, 600,000 GPU-hrs/moGrowth
- 600,000+ GPU-hrs/moData Center
Run a 14-day Scale evaluation before you buy
Request an evaluation code for your own logs. A verified workplace organization activates up to 50 diagnoses a day. Start with one resolved incident and compare the result with the confirmed cause and action. No card and no automatic subscription.
Want a guided replay with agreed success criteria? Scope an incident evaluation.
Review the Terms of Service, Privacy Policy and Refund policy. The 14-day Scale evaluation requires no card and creates no automatic subscription.
Free
Free forever
- T0 (anonymous or negatively classified): 3/day. T1 (identified): 10/day. T2 (current verified organization proof): 50/day. T3 (active trial): the persisted trial cap, normally 50/day. T4 (paid): unlimited.
- Paste logs in the web UI, nothing to install
- 10,800+ deterministic matcher rules
- AI fallback for novel errors within the current daily allowance
- Evidence-ranked root cause + concrete next action, not an essay
- Cost optimization advice on every diagnosis (evidence-based right-sizing, spot vs on-demand)
- Inference serving too: vLLM, SGLang, TensorRT-LLM, LMDeploy, TGI, Triton, engine deaths, KV-cache exhaustion, disaggregated prefill/decode
Team
Billed annually · $4,980/yr · 2 months free
25k GPU-hours included · $0.06/GPU-hr overage
- Everything in Free, plus:
- Unlimited seats
- Diagnose jobs up to 128 GPUs
- Lightweight live fleet status + Fleet Readiness for up to 128 GPUs
- Fleet change correlation: compares failed hosts with healthy controls to show what changed
- All 16,400+ failure signatures + AI fallback for novel errors
- Cross-rank cascade analysis: isolates the initiating rank when evidence supports it
- NCCL hang culprit-rank localizer
Scale
Billed annually · $29,940/yr · 2 months free
150k GPU-hours included · $0.04/GPU-hr overage
- Everything in Team, plus:
- Advanced proactive monitoring up to 1,024 GPUs: prediction, telemetry history, escalation, and scheduler controls
- In-VPC agent option & self-hosted OpenAlex mirrors: logs and research queries never leave your cluster
- Governed remediation dry runs with exact human-run commands and rollback guidance
- L1 to L2 to L3 auto-escalation on recurring incidents
- Silent data corruption (SDC) detection
- Straggler + gray failure detection
- DCGM thermal peer-comparison (micro-stragglers)
Growth
Billed annually · $54,960/yr · 2 months free
600k GPU-hours included · $0.02/GPU-hr overage
- Everything in Scale, plus:
- Monitor up to 4,096 GPUs
- Costs less than Scale above ~202,500 GPU-hours/month (4x the included hours, half the overage rate)
- Priority support: 1-business-day P1 response
Price protection: existing subscribers keep their rate and included GPU-hours when list prices rise.
Every plan includes one-click connectors. Slack, PagerDuty, W&B, MLflow, TensorBoard, never metered.
Data Center
$12,500/ monthBilled annually · $150,000/yr · 2 months free2M GPU-hours included · $0.015/GPU-hr overage · volume & per-node pricing
Designed for fleets up to 16,384 GPUs · multi-tenant · white-label / OEM. Volume per-GPU, or per-node pricing for GPU-cloud providers who bill their own customers by the node. Onboarded through a scoped pilot, then scaled to your full fleet.
- Designed for fleets up to 16,384 GPUs. Beyond by pilot, multi-tenant
- Mass-crash coalescing: one fabric event → one diagnosis, one page
- White-label / OEM diagnosis for your customers
- Predictive failure scoring on every diagnosis; remote execution by controlled-pilot qualification
- Checkpoint integrity analysis + launcher-specific rollback/resume runbooks
- SLURM and Kubernetes integration; Ray failure diagnosis
- BYOK · EU region (planned) · air-gapped on-host diagnosis engine (hosted control plane excluded)
- HIPAA BAA · Live status page · availability terms on contract · dedicated CSM
On-premise in-VPC agent on Scale+ · Upgrade or cancel anytime
Plan changes take effect immediately with prorated billing. On downgrades, the unused portion credits to your next invoice.
Need something custom? Talk to sales.
From diagnosis to autonomy
Diagnosis closes the loop on understanding. The autonomy layer closes the loop on recovery. See the autonomy roadmap for what's shipped, in progress, and planned.
Predictive failure scoring
Telemetry models flag deteriorating GPUs and nodes before the crash, so jobs migrate instead of dying.
Automated node cordoning
A bad GPU is marked unhealthy and removed from the scheduler automatically.
Automated checkpoint rollback
On failure: find the last good checkpoint, validate it, resume training. No human in the loop.
Auto-remediation engine
Detect → cordon → roll back → resume. Autonomous recovery that closes the loop end to end.
Track delivery dates on the product roadmap →
Security & compliance posture
We label compliance honestly: SOC 2 Type II is planned, not claimed. Read the Trust Center →
Frequently asked questions
Your next failure is already scheduled.
The only question is whether it costs you twelve hours of grep, or one paste into Denpex.