A 32-node DDP job ends in an NCCL timeout after one rank exhausts memory. The useful path is to identify the first failed rank, prove the memory growth, and change the responsible configuration before investigating the fabric.
Training crashed? Find what failed first.
Paste your training log. Get the likely root cause, the evidence behind it, and the next safe steps. See which failure started the cascade.
3 free cloud diagnoses daily. No signup, credit card, or installation.
Need help now? Start with your log. Evaluating for your team? Compare a diagnosis with your own confirmed remedy, then prove notifications and report access on one connected workload.
Example diagnosis
Animated demo. A 64 GPU Llama 3 70B fine-tuning run crashes. Denpex reads the logs and reports: symptom, 64 NCCL collective timeouts that were only collateral; cause, rank 42 on node-05, GPU 8A:00, Xid 48 double-bit ECC error; class, hardware, route to infra rather than ML. It illustrates reviewing node isolation and resuming from a verified checkpoint if workload compatibility and restart safety are established, then alerts the on-call engineer with the recommendation attached. This is an illustration, not a customer recovery result.
Built for
Frontier LLM training
Evaluate cross-rank cascade analysis with your workload. Numerical and checksum evidence can expose corruption indicators; checkpoint resume requires verified integrity and compatibility. Large-cluster acceptance is customer specific.
FSDP / DDP production
PyTorch DDP and FSDP at scale. Diagnose the originating rank, the failure class, and the evidence-backed next action.
GPU clouds and neoclouds
White-label diagnosis for your customers. Offer Denpex under your brand, on the clusters you already operate.
Healthcare and life sciences
HIPAA BAA on Data Center. Configurable PHI masking on the agent, with audit logs of every diagnosis.
Evaluate Denpex with your incidents. Scope a pilot →
Built for distributed ML training.
From frontier LLM training to FSDP fine-tunes to GPU cloud operations. Denpex fits the way your team already runs.
Frontier LLM training
Evaluate cross-rank cascade analysis with your workload. Numerical and checksum evidence can expose corruption indicators; checkpoint resume requires verified integrity and compatibility. Large-cluster acceptance is customer specific.
FSDP / DDP production
PyTorch DDP and FSDP at scale. Diagnose the originating rank, the failure class, and the evidence-backed next action.
GPU clouds and neoclouds
White-label diagnosis for your customers. Offer Denpex under your brand, on the clusters you already operate.
Healthcare and life sciences
HIPAA BAA on Data Center. Configurable PHI masking on the agent, with audit logs of every diagnosis.
Supported Ecosystems
Frameworks
Schedulers
Monitoring
Who uses Denpex?
Primary Operators
Training Infrastructure Engineers, ML Infrastructure Engineers, AI Platform Engineers, MLOps Engineers, Distributed Systems Engineers, GPU Cluster Engineers, SREs, and Platform Reliability Engineers.
Secondary Beneficiaries
AI Researchers, Machine Learning Engineers, Deep Learning Engineers, Data Scientists, and Performance Engineers.
Platform Buyers
Director of AI Infrastructure, Head of ML Platform, VP of Engineering, CTO, and Director of Research Infrastructure.
Ideal Organizations
AI startups training foundation models, Enterprise AI teams, GPU cloud providers, Managed ML platforms, Universities, National laboratories (HPC), Autonomous vehicle companies, and Defense contractors.
Your logs say “NCCL timeout.” The timeline says rank 17.
Distributed timeline reconstruction orders every event across every rank, clock-drift corrected, so the cascade reads in causal order, not log order. The watchdog is the last thing that happened, never the first.
The complete reliability surface
Diagnosis is the entry point. The platform covers the whole failure lifecycle before launch, during training, and after the fix ships.
10,800+ deterministic signatures plus retrieval over 700+ published Failure Encyclopedia entries, with deeper fallback analysis for anything novel. Every diagnosis ends in a concrete next action.
Root cause analysis
The best-supported initiating fault, separated from the symptom that woke you up
First-failed-rank detection
The initiating rank when the submitted evidence supports attribution; unknown otherwise
Distributed rank correlation
Cross-rank telemetry stitched into one causal picture
Cascade failure analysis
How one bad GPU took 63 healthy ranks down with it
NCCL timeout diagnosis
The initiator behind the watchdog's generic timeout
CUDA OOM diagnosis + memory attribution
The tensor or layer when allocator or profiler evidence identifies it; otherwise the next discriminator
Memory fragmentation diagnosis
Reserved-but-unallocated signatures, allocator-level fixes
Gradient explosion diagnosis
Norm spikes traced back to layer and step
NaN loss diagnosis
The propagation source, not just the first poisoned batch
Weight divergence diagnosis
Drift measured against your own healthy baselines
Silent hang diagnosis
Heartbeat detection for jobs that die without a stack trace
Device assert diagnosis
Device-side asserts mapped to the offending operation
Checkpoint corruption diagnosis
Torn writes and truncated shards caught before resume
Import error diagnosis
Environment faults separated from training faults
Version mismatch diagnosis
PyTorch x CUDA x cuDNN conflicts flagged precisely
Disk full diagnosis
Storage exhaustion before it masquerades as a framework crash
AI fallback analysis
Unknown failures get deep analysis on masked excerpts
Prescriptive fixes
Copy-paste resolution paths, verified against the failure class
Resume checkpoint recommendations
The last verified-good step to restart from
Hardware vs software classification
Infra issue or ML issue, instantly, so the right team moves
Looking for the full encyclopedia? Browse 700+ published entries →
From a failed job to a next step supported by evidence.
When your connected training job fails, Denpex collects the evidence and starts investigating. Get a specific next action, the facts supporting it, and checks to determine whether your workload recovered.
In-VPC agent
Single Python file, stdlib only, no root, no kernel module. Wraps your training command, heartbeats every 120 s. PII/PHI masking runs client-side, before any byte leaves your cluster, on by default. Set DENPEX_PRIVACY=strict and the agent ships only anonymized failure signatures; raw logs never leave.
Deterministic engine
10,800+ deterministic regex signatures plus IDF-weighted retrieval over 700+ published Failure Encyclopedia entries, with clock-drift-corrected timeline reconstruction, do the work deterministically. The AI fallback only sees masked excerpts of novel failures, and your logs are never used to train anything.
Routed resolution
Ownership mapping sends one correlated incident, evidence-ranked cause, classification, and concrete next action to the engineer who owns the job, on Slack, PagerDuty, SMS or webhook. Hardware issues route to infra; ML issues route to research.
What your failures cost. And what you get back.
Estimate the GPU-hours and engineering time your team loses to undiagnosed failures, and the payback on each plan.
The kind of failure this catches
Illustrative incident paths based on documented failure modes in the Denpex encyclopedia. They are not customer quotes or claims about a specific engagement.
An FSDP fine-tune fails only when a particular sequence shape reaches a corrupted sample. Reproducing the same sample on a known-good node separates the data path from an unnecessary hardware escalation.
A distributed job reports a collective timeout after Rank 47 hits a CUDA OOM. A shared evidence trail lets the model and infrastructure teams work from the initiating failure instead of treating the final timeout as the cause.
Priced against your GPU bill, not your seat count.
A single failure on a 64-GPU cluster wastes hours of compute and an afternoon of engineering time. Every plan starts free, no credit card needed. Annual plans save 2 months.
Not sure which plan? Match your monthly GPU-hours.
- 0, 25,000 GPU-hrs/moTeam
- 25,000, 150,000 GPU-hrs/moScale
- 150,000, 600,000 GPU-hrs/moGrowth
- 600,000+ GPU-hrs/moData Center
Run a 30-day Team evaluation before you buy
Request an evaluation code for your own logs. A verified workplace organization activates up to 50 diagnoses a day. Start with one resolved incident and compare the result with the confirmed cause and action. No card and no automatic subscription.
Want a guided replay with agreed success criteria? Scope an incident evaluation.
Review the Terms of Service, Privacy Policy and Refund policy. The 30-day Team evaluation requires no card and creates no automatic subscription.
Free
Free forever
- 3 diagnoses daily without signup, 10 for identified accounts. Verified organizations can qualify for 50 daily; your current allowance is shown before diagnosis. Trial limits are shown in your account. Paid plans include unlimited diagnoses.
- Paste logs in the web UI, nothing to install
- 10,800+ deterministic matcher rules
- AI fallback for novel errors within the current daily allowance
- Evidence-ranked root cause + concrete next action, not an essay
- Cost optimization recommendations when workload and pricing evidence support them
- Inference serving too: vLLM, SGLang, TensorRT-LLM, LMDeploy, TGI, Triton, engine deaths, KV-cache exhaustion, disaggregated prefill/decode
Team
Billed annually · $4,980/yr · 2 months free
25k GPU-hours included · $0.06/GPU-hr overage
- Everything in Free, plus:
- Unlimited seats
- Diagnose jobs up to 128 GPUs
- Lightweight live fleet status + Fleet Readiness for up to 128 GPUs
- Fleet change correlation: compares failed hosts with healthy controls to show what changed
- All 16,400+ failure signatures + AI fallback for novel errors
- Cross-rank cascade analysis: isolates the initiating rank when evidence supports it
- NCCL hang culprit-rank localizer
Scale
Billed annually · $29,940/yr · 2 months free
150k GPU-hours included · $0.04/GPU-hr overage
- Everything in Team, plus:
- Advanced proactive monitoring up to 1,024 GPUs: prediction, telemetry history, escalation, and scheduler controls
- Configured in-VPC local diagnosis and self-hosted research mirrors; hosted control plane and cloud alerts are separate
- Governed remediation dry runs with exact human-run commands and rollback guidance
- L1 to L2 to L3 auto-escalation on recurring incidents
- Silent data corruption indicators from supplied checksums and numerical telemetry; confirm with a trusted control
- Straggler + gray failure detection
- DCGM thermal peer-comparison (micro-stragglers)
Growth
Billed annually · $54,960/yr · 2 months free
600k GPU-hours included · $0.02/GPU-hr overage
- Everything in Scale, plus:
- Monitor up to 4,096 GPUs
- Costs less than Scale above ~202,500 GPU-hours/month (4x the included hours, half the overage rate)
- Priority support: 1-business-day P1 response
Price protection: existing subscribers keep their rate and included GPU-hours when list prices rise.
Every plan includes connectors with setup guides. Slack, PagerDuty, W&B, MLflow, TensorBoard, never metered.
Data Center
$12,500/ monthBilled annually · $150,000/yr · 2 months free2M GPU-hours included · $0.015/GPU-hr overage · volume & per-node pricing
Designed for fleets up to 16,384 GPUs · multi-tenant · white-label / OEM. Volume per-GPU, or per-node pricing for GPU-cloud providers who bill their own customers by the node. Onboarded through a scoped pilot, then scaled to your full fleet.
- Designed for fleets up to 16,384 GPUs. Beyond by pilot, multi-tenant
- Mass-crash coalescing for correlated incidents; validate grouping and delivery in your pilot
- White-label / OEM diagnosis for your customers
- Predictive failure scoring on every diagnosis; remote execution by controlled-pilot qualification
- Checkpoint integrity analysis + launcher-specific rollback/resume runbooks
- SLURM and Kubernetes integration; Ray failure diagnosis
- BYOK · EU region (planned) · air-gapped on-host diagnosis engine (hosted control plane excluded)
- HIPAA BAA · Live status page · availability terms on contract · dedicated CSM
On-premise in-VPC agent on Scale+ · Upgrade or cancel anytime
Plan changes take effect immediately with prorated billing. On downgrades, the unused portion credits to your next invoice.
Need something custom? Talk to sales.
From diagnosis to autonomy
Connected jobs can submit evidence and start an investigation. Recommendations require workload verification; remote execution is restricted to qualified, authorized pilots. See the autonomy roadmap for what's shipped, in progress, and planned.
Predictive failure scoring
Scoring available; migration plannedEvidence-based indicators help prioritize investigation. Predicting every failure or migrating jobs automatically is not a generally available capability.
Automated node cordoning
Controlled pilotQualified scheduler actions require a scoped connection, target and prerequisites. General unattended node removal is not included in ordinary diagnosis.
Automated checkpoint rollback
Controlled pilotCheckpoint assessment and resume guidance are available. Execution requires validated checkpoint integrity, workload compatibility and explicit authorization.
Auto-remediation engine
Controlled pilotRemediation qualification and dry runs support operator review. Arbitrary autonomous cluster repair is not offered as a guaranteed production capability.
Track delivery dates on the product roadmap →
Security & compliance posture
We label compliance honestly: SOC 2 Type II is planned, not claimed. Read the Trust Center →
Frequently asked questions
Your next failure is already scheduled.
Evaluate one incident from your environment. Compare the recommendation with the confirmed remedy, then prove a connected workload's notification and report path.