Skip to content
GPU training diagnosis, no signup

Training crashed? Find what failed first.

Paste your training log. Get the likely root cause, the evidence behind it, and the next safe steps. See which failure started the cascade.

3 free cloud diagnoses daily. No signup, credit card, or installation.

Need help now? Start with your log. Evaluating for your team? Compare a diagnosis with your own confirmed remedy, then prove notifications and report access on one connected workload.

Example diagnosis

Animated demo. A 64 GPU Llama 3 70B fine-tuning run crashes. Denpex reads the logs and reports: symptom, 64 NCCL collective timeouts that were only collateral; cause, rank 42 on node-05, GPU 8A:00, Xid 48 double-bit ECC error; class, hardware, route to infra rather than ML. It illustrates reviewing node isolation and resuming from a verified checkpoint if workload compatibility and restart safety are established, then alerts the on-call engineer with the recommendation attached. This is an illustration, not a customer recovery result.

llama3-70b-finetune-run-47
CrashDiagnoseFixAlert

Built for

Frontier LLM training

Evaluate cross-rank cascade analysis with your workload. Numerical and checksum evidence can expose corruption indicators; checkpoint resume requires verified integrity and compatibility. Large-cluster acceptance is customer specific.

FSDP / DDP production

PyTorch DDP and FSDP at scale. Diagnose the originating rank, the failure class, and the evidence-backed next action.

GPU clouds and neoclouds

White-label diagnosis for your customers. Offer Denpex under your brand, on the clusters you already operate.

Healthcare and life sciences

HIPAA BAA on Data Center. Configurable PHI masking on the agent, with audit logs of every diagnosis.

Evaluate Denpex with your incidents. Scope a pilot →

Built for distributed ML training.

From frontier LLM training to FSDP fine-tunes to GPU cloud operations. Denpex fits the way your team already runs.

Frontier LLM training

Evaluate cross-rank cascade analysis with your workload. Numerical and checksum evidence can expose corruption indicators; checkpoint resume requires verified integrity and compatibility. Large-cluster acceptance is customer specific.

FSDP / DDP production

PyTorch DDP and FSDP at scale. Diagnose the originating rank, the failure class, and the evidence-backed next action.

GPU clouds and neoclouds

White-label diagnosis for your customers. Offer Denpex under your brand, on the clusters you already operate.

Healthcare and life sciences

HIPAA BAA on Data Center. Configurable PHI masking on the agent, with audit logs of every diagnosis.

Supported Ecosystems

Frameworks

PyTorchDeepSpeedHorovodMegatron-LMRayHugging Face Accelerate

Schedulers

SLURMKubernetes

Monitoring

Weights & BiasesMLflowTensorBoard

Who uses Denpex?

Primary Operators

Training Infrastructure Engineers, ML Infrastructure Engineers, AI Platform Engineers, MLOps Engineers, Distributed Systems Engineers, GPU Cluster Engineers, SREs, and Platform Reliability Engineers.

Secondary Beneficiaries

AI Researchers, Machine Learning Engineers, Deep Learning Engineers, Data Scientists, and Performance Engineers.

Platform Buyers

Director of AI Infrastructure, Head of ML Platform, VP of Engineering, CTO, and Director of Research Infrastructure.

Ideal Organizations

AI startups training foundation models, Enterprise AI teams, GPU cloud providers, Managed ML platforms, Universities, National laboratories (HPC), Autonomous vehicle companies, and Defense contractors.

Your logs say “NCCL timeout.” The timeline says rank 17.

Distributed timeline reconstruction orders every event across every rank, clock-drift corrected, so the cascade reads in causal order, not log order. The watchdog is the last thing that happened, never the first.

t+0 msrank 17ECC uncorrectable error: the true initiatorROOT CAUSE
t+340 msrank 4stalls waiting on the wedged collective
t+600 sall ranksNCCL watchdog times out, the only line your logs showed you

The complete reliability surface

Diagnosis is the entry point. The platform covers the whole failure lifecycle before launch, during training, and after the fix ships.

10,800+ deterministic signatures plus retrieval over 700+ published Failure Encyclopedia entries, with deeper fallback analysis for anything novel. Every diagnosis ends in a concrete next action.

Root cause analysis

The best-supported initiating fault, separated from the symptom that woke you up

First-failed-rank detection

The initiating rank when the submitted evidence supports attribution; unknown otherwise

Distributed rank correlation

Cross-rank telemetry stitched into one causal picture

Cascade failure analysis

How one bad GPU took 63 healthy ranks down with it

NCCL timeout diagnosis

The initiator behind the watchdog's generic timeout

CUDA OOM diagnosis + memory attribution

The tensor or layer when allocator or profiler evidence identifies it; otherwise the next discriminator

Memory fragmentation diagnosis

Reserved-but-unallocated signatures, allocator-level fixes

Gradient explosion diagnosis

Norm spikes traced back to layer and step

NaN loss diagnosis

The propagation source, not just the first poisoned batch

Weight divergence diagnosis

Drift measured against your own healthy baselines

Silent hang diagnosis

Heartbeat detection for jobs that die without a stack trace

Device assert diagnosis

Device-side asserts mapped to the offending operation

Checkpoint corruption diagnosis

Torn writes and truncated shards caught before resume

Import error diagnosis

Environment faults separated from training faults

Version mismatch diagnosis

PyTorch x CUDA x cuDNN conflicts flagged precisely

Disk full diagnosis

Storage exhaustion before it masquerades as a framework crash

AI fallback analysis

Unknown failures get deep analysis on masked excerpts

Prescriptive fixes

Copy-paste resolution paths, verified against the failure class

Resume checkpoint recommendations

The last verified-good step to restart from

Hardware vs software classification

Infra issue or ML issue, instantly, so the right team moves

Looking for the full encyclopedia? Browse 700+ published entries →

From a failed job to a next step supported by evidence.

When your connected training job fails, Denpex collects the evidence and starts investigating. Get a specific next action, the facts supporting it, and checks to determine whether your workload recovered.

Request the full architecture brief
01 · your boundary

In-VPC agent

Single Python file, stdlib only, no root, no kernel module. Wraps your training command, heartbeats every 120 s. PII/PHI masking runs client-side, before any byte leaves your cluster, on by default. Set DENPEX_PRIVACY=strict and the agent ships only anonymized failure signatures; raw logs never leave.

02 · pattern-first, AI-last

Deterministic engine

10,800+ deterministic regex signatures plus IDF-weighted retrieval over 700+ published Failure Encyclopedia entries, with clock-drift-corrected timeline reconstruction, do the work deterministically. The AI fallback only sees masked excerpts of novel failures, and your logs are never used to train anything.

03 · one incident, one owner

Routed resolution

Ownership mapping sends one correlated incident, evidence-ranked cause, classification, and concrete next action to the engineer who owns the job, on Slack, PagerDuty, SMS or webhook. Hardware issues route to infra; ML issues route to research.

What your failures cost. And what you get back.

Estimate the GPU-hours and engineering time your team loses to undiagnosed failures, and the payback on each plan.

The kind of failure this catches

Illustrative incident paths based on documented failure modes in the Denpex encyclopedia. They are not customer quotes or claims about a specific engagement.

OOM masked as NCCL

A 32-node DDP job ends in an NCCL timeout after one rank exhausts memory. The useful path is to identify the first failed rank, prove the memory growth, and change the responsible configuration before investigating the fabric.

Illustrative scenario32 node DDP cluster
Dataloader, not hardware

An FSDP fine-tune fails only when a particular sequence shape reaches a corrupted sample. Reproducing the same sample on a known-good node separates the data path from an unnecessary hardware escalation.

Illustrative scenarioFSDP fine tuning
Ends the 2am blame game

A distributed job reports a collective timeout after Rank 47 hits a CUDA OOM. A shared evidence trail lets the model and infrastructure teams work from the initiating failure instead of treating the final timeout as the cause.

Illustrative scenarioShared training infrastructure

Priced against your GPU bill, not your seat count.

A single failure on a 64-GPU cluster wastes hours of compute and an afternoon of engineering time. Every plan starts free, no credit card needed. Annual plans save 2 months.

Annual2 months free

Not sure which plan? Match your monthly GPU-hours.

  • 0, 25,000 GPU-hrs/moTeam
  • 25,000, 150,000 GPU-hrs/moScale
  • 150,000, 600,000 GPU-hrs/moGrowth
  • 600,000+ GPU-hrs/moData Center

Run a 30-day Team evaluation before you buy

Request an evaluation code for your own logs. A verified workplace organization activates up to 50 diagnoses a day. Start with one resolved incident and compare the result with the confirmed cause and action. No card and no automatic subscription.

Use your company email (e.g. alex@yourcompany.com). Established company domains can activate automatically after email and domain checks, regardless of mail host. Uncertain domains can request review from support@denpex.com. Free local diagnoses remain available without an account.

Want a guided replay with agreed success criteria? Scope an incident evaluation.

Review the Terms of Service, Privacy Policy and Refund policy. The 30-day Team evaluation requires no card and creates no automatic subscription.

Free

$0/ month

Free forever

  • 3 diagnoses daily without signup, 10 for identified accounts. Verified organizations can qualify for 50 daily; your current allowance is shown before diagnosis. Trial limits are shown in your account. Paid plans include unlimited diagnoses.
  • Paste logs in the web UI, nothing to install
  • 10,800+ deterministic matcher rules
  • AI fallback for novel errors within the current daily allowance
  • Evidence-ranked root cause + concrete next action, not an essay
  • Cost optimization recommendations when workload and pricing evidence support them
  • Inference serving too: vLLM, SGLang, TensorRT-LLM, LMDeploy, TGI, Triton, engine deaths, KV-cache exhaustion, disaggregated prefill/decode
Most popular

Team

$415/ month

Billed annually · $4,980/yr · 2 months free

25k GPU-hours included · $0.06/GPU-hr overage

  • Everything in Free, plus:
  • Unlimited seats
  • Diagnose jobs up to 128 GPUs
  • Lightweight live fleet status + Fleet Readiness for up to 128 GPUs
  • Fleet change correlation: compares failed hosts with healthy controls to show what changed
  • All 16,400+ failure signatures + AI fallback for novel errors
  • Cross-rank cascade analysis: isolates the initiating rank when evidence supports it
  • NCCL hang culprit-rank localizer

Scale

$2,495/ month

Billed annually · $29,940/yr · 2 months free

150k GPU-hours included · $0.04/GPU-hr overage

  • Everything in Team, plus:
  • Advanced proactive monitoring up to 1,024 GPUs: prediction, telemetry history, escalation, and scheduler controls
  • Configured in-VPC local diagnosis and self-hosted research mirrors; hosted control plane and cloud alerts are separate
  • Governed remediation dry runs with exact human-run commands and rollback guidance
  • L1 to L2 to L3 auto-escalation on recurring incidents
  • Silent data corruption indicators from supplied checksums and numerical telemetry; confirm with a trusted control
  • Straggler + gray failure detection
  • DCGM thermal peer-comparison (micro-stragglers)

Growth

$4,580/ month

Billed annually · $54,960/yr · 2 months free

600k GPU-hours included · $0.02/GPU-hr overage

  • Everything in Scale, plus:
  • Monitor up to 4,096 GPUs
  • Costs less than Scale above ~202,500 GPU-hours/month (4x the included hours, half the overage rate)
  • Priority support: 1-business-day P1 response

Price protection: existing subscribers keep their rate and included GPU-hours when list prices rise.

Every plan includes connectors with setup guides. Slack, PagerDuty, W&B, MLflow, TensorBoard, never metered.

Data Center

$12,500/ monthBilled annually · $150,000/yr · 2 months free

2M GPU-hours included · $0.015/GPU-hr overage · volume & per-node pricing

Designed for fleets up to 16,384 GPUs · multi-tenant · white-label / OEM. Volume per-GPU, or per-node pricing for GPU-cloud providers who bill their own customers by the node. Onboarded through a scoped pilot, then scaled to your full fleet.

  • Designed for fleets up to 16,384 GPUs. Beyond by pilot, multi-tenant
  • Mass-crash coalescing for correlated incidents; validate grouping and delivery in your pilot
  • White-label / OEM diagnosis for your customers
  • Predictive failure scoring on every diagnosis; remote execution by controlled-pilot qualification
  • Checkpoint integrity analysis + launcher-specific rollback/resume runbooks
  • SLURM and Kubernetes integration; Ray failure diagnosis
  • BYOK · EU region (planned) · air-gapped on-host diagnosis engine (hosted control plane excluded)
  • HIPAA BAA · Live status page · availability terms on contract · dedicated CSM
Book a demo

On-premise in-VPC agent on Scale+ · Upgrade or cancel anytime

Plan changes take effect immediately with prorated billing. On downgrades, the unused portion credits to your next invoice.

Need something custom? Talk to sales.

From diagnosis to autonomy

Connected jobs can submit evidence and start an investigation. Recommendations require workload verification; remote execution is restricted to qualified, authorized pilots. See the autonomy roadmap for what's shipped, in progress, and planned.

Predictive failure scoring

Scoring available; migration planned

Evidence-based indicators help prioritize investigation. Predicting every failure or migrating jobs automatically is not a generally available capability.

Automated node cordoning

Controlled pilot

Qualified scheduler actions require a scoped connection, target and prerequisites. General unattended node removal is not included in ordinary diagnosis.

Automated checkpoint rollback

Controlled pilot

Checkpoint assessment and resume guidance are available. Execution requires validated checkpoint integrity, workload compatibility and explicit authorization.

Auto-remediation engine

Controlled pilot

Remediation qualification and dry runs support operator review. Arbitrary autonomous cluster repair is not offered as a guaranteed production capability.

Track delivery dates on the product roadmap →

Security & compliance posture

Logs deleted after diagnosis on Free/TeamPII / PHI masking before egressIn-VPC agent on Scale and Data CenterRetention + purge controlsTeam roles (owner / admin / member) · SSO (SAML/OIDC) planned Q4 2026GDPR DPA · HIPAA BAA on Data Center · SOC 2 Type II planned

We label compliance honestly: SOC 2 Type II is planned, not claimed. Read the Trust Center →

Frequently asked questions

Your next failure is already scheduled.

Evaluate one incident from your environment. Compare the recommendation with the confirmed remedy, then prove a connected workload's notification and report path.