A single failure on a 64-GPU cluster wastes hours of compute and an afternoon of engineering time. Every plan starts free, no credit card needed. Annual plans save 2 months.
Free forever
Billed annually · $4,980/yr · 2 months free
25k GPU-hours included · $0.06/GPU-hr overage
Billed annually · $29,940/yr · 2 months free
150k GPU-hours included · $0.04/GPU-hr overage
Billed annually · $54,960/yr · 2 months free
600k GPU-hours included · $0.02/GPU-hr overage
Price protection: existing subscribers keep their rate and included GPU-hours when list prices rise.
Every plan includes one-click connectors. Slack, PagerDuty, W&B, MLflow, TensorBoard, never metered.
2M GPU-hours included · $0.015/GPU-hr overage · volume & per-node pricing
Designed for fleets up to 16,384 GPUs · multi-tenant · white-label / OEM. Volume per-GPU, or per-node pricing for GPU-cloud providers who bill their own customers by the node. Onboarded through a scoped pilot, then scaled to your full fleet.
On-premise in-VPC agent on Scale+ · Upgrade or cancel anytime
Plan changes take effect immediately with prorated billing. On downgrades, the unused portion credits to your next invoice.
Need something custom? Talk to sales.
Calculate the hidden infrastructure waste from fail-slow incidents.
Pipeline-parallel training runs at the speed of the slowest node. A 30% slower node wastes 30% of the entire cluster's compute. 100% represents a full stall or rollback.
Wasted compute every time this straggler pattern hits the cluster.
Denpex detects thermal and memory stragglers within 30 seconds of degradation, auto-fencing the node before the pipeline bubble expands.
Concrete controls, not vague promises. We label compliance honestly: SOC 2 Type II is planned, not claimed.
Client-side masking runs on the agent before any log is transmitted. Default patterns catch emails, SSNs, phone numbers, credit cards, and common PHI (MRN, NPI). Add your own patterns. Raw PII/PHI never leaves your cluster.
On Free and Team, raw logs are processed in memory and never written to durable storage. We retain anonymized failure signatures and resolution metadata only, never raw lines. An in-VPC agent (no log egress at all) ships on Scale and Data Center.
TLS 1.3 in transit, AES-256 at rest. Encryption keys are customer-managed (BYOK) on Data Center via your KMS.
SOC 2 Type II planned. GDPR DPA available, HIPAA BAA available on Data Center. Sub-processor list and data flow on the Trust Center.
SSO via Google, Discord, GitHub, and Microsoft. Enterprise SAML/OIDC SSO is on our roadmap. Role-based access (owner, admin, member, viewer) on Team+. 30-day audit log of all billing and team changes.
Cloud (default), single-tenant on AWS / Azure / GCP, or fully air-gapped on your hardware on Data Center. White-label / OEM available for GPU clouds.