Skip to content
Public dataset · updated with the encyclopedia

GPU Reliability Index

Teams commit eight figures to GPU fleets with less reliability data than a used-car buyer gets. This page is the antidote, in two honestly-separated layers: published field studies (what large clusters actually measured, with sources), and the documented failure surface per GPU family, interconnect and provider stack from the 559 classified signatures in our public Failure Encyclopedia.

What the field actually measured

Published third-party reliability measurements, reported as-published with sources. Denpex reprints these; it did not produce them.

Meta — The Llama 3 Herd of Models (infrastructure section), 2024

16,384× NVIDIA H100 80GB (SXM), Meta production training · 54-day Llama-3 405B pre-training run

  • Unexpected interruptions: 419 in 54 days (~1 every 3 hours cluster-wide)
  • Hardware share of interruptions: ~78%
  • Faulty GPUs (incl. NVLink): 30.1% of unexpected interruptions
  • HBM3 memory faults: 17.2% of unexpected interruptions

Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs (NCSA Delta / DeltaAI)

1,056 GPUs (A100 + H100), 11.7M GPU-hours over 2.5 years · 2.5 years of production operation

  • H100 vs A100 memory errors: H100 shows ~3.2× LOWER per-GPU mean-time-between-(memory)-errors than A100 — HBM3 error handling has not kept pace with capacity
  • H100 core hardware: H100 shows improved resilience vs A100 on critical (non-memory) hardware components
  • Capacity planning: ~5% GPU overprovisioning projected necessary at scale to absorb failures

Understanding GPU Memory Corruption at Extreme Scale: The Summit Case Study (ORNL)

Oak Ridge Summit — 27,648× V100 · multi-year production study

  • Scope: Large-scale characterization of GPU memory corruption (SBE/DBE) behavior and its job-level impact on V100 at extreme scale

Understanding the Landscape of Ampere GPU Memory Errors (2025)

Multi-cluster A100 study · multi-year

  • Scope: Cross-cluster analysis of A100 memory-error rates, spatial locality of faults, and row-remapping effectiveness

What we deliberately do NOT publish yet: a Denpex-fleet MTBF table. Cross-tenant incident telemetry only began accumulating in July 2026, and an index computed from weeks of data would be noise sold as signal. When the corpus crosses 90 days and the cross-tenant disclosure ships, this section lights up with real numbers and full denominators.

The documented failure surface

Counts of DISTINCT documented failure signatures in the public Denpex Failure Encyclopedia (559 published classes). These measure how many documented, fixable ways each slice of the stack fails — a coverage/complexity map for procurement diligence, not a field failure rate. Every count links back to entries any engineer can read in the Failure Encyclopedia — this is auditable, not vibes. 61 of 559 signatures are hardware-class; 13 NVIDIA Xid codes are covered (13, 31, 43, 48, 61, 62, 63, 64, 69, 74, 79, 94, 119).

By GPU family

How many DISTINCT documented ways training fails on/around each GPU family — a complexity map for diligence, not a field failure rate.

SliceDocumented failure signaturesCritical-severityWhere they concentrate
H100 (SXM/PCIe)103Reliability (2) · Hardware (2) · Communication (2)
A10041Reliability (2) · Hardware (2)
H20021Reliability (1) · Communication (1)
B200 / GB20021Reliability (1) · Communication (1)

By interconnect

NVLink/NVSwitch vs InfiniBand vs RoCE vs EFA vs PCIe — where the fabric bites.

SliceDocumented failure signaturesCritical-severityWhere they concentrate
PCIe228Hardware (14) · Communication (2) · Memory (1)
NVLink / NVSwitch186Hardware (8) · Communication (4) · Fail-Slow (1)
InfiniBand173Communication (10) · Network (2) · Distributed Communication (1)
RoCE / Ethernet72Communication (4) · Network (2) · Hardware (1)
AWS EFA31Communication (3)

By provider stack

Failure signatures that reference each provider's stack (EFA, GKE, AKS, Slurm/on-prem, ...). Higher ≠ worse cloud — it can mean more usage and better-documented failure modes. Read with the studies above.

SliceDocumented failure signaturesCritical-severityWhere they concentrate
Kubernetes (any cloud)100Infrastructure (5) · Communication (2) · Reliability (1)
On-prem / HPC (Slurm etc.)81Infrastructure (4) · Distributed Training (1) · Environment (1)
AWS71Communication (3) · Infrastructure (2) · Network (1)
Azure30Infrastructure (3)
Google Cloud20Infrastructure (1) · Network (1)
CoreWeave20Reliability (1) · Infrastructure (1)
RunPod20Infrastructure (2)
Lambda11Infrastructure (1)

Take the dataset with you

Machine-readable JSON (free, no key): GET https://api.denpex.com/api/reliability-index. Cite it in your procurement memo; the sources above carry the field numbers.

Browse the 559-class Failure Encyclopedia