Teams commit eight figures to GPU fleets with less reliability data than a used-car buyer gets. This page is the antidote, in two honestly-separated layers: published field studies (what large clusters actually measured, with sources), and the documented failure surface per GPU family, interconnect and provider stack from the 559 classified signatures in our public Failure Encyclopedia.
What the field actually measured
Published third-party reliability measurements, reported as-published with sources. Denpex reprints these; it did not produce them.
Meta. The Llama 3 Herd of Models (infrastructure section), 2024
16,384× NVIDIA H100 80GB (SXM), Meta production training · 54-day Llama-3 405B pre-training run
Unexpected interruptions: 419 in 54 days (~1 every 3 hours cluster-wide)
Hardware share of interruptions: ~78%
Faulty GPUs (incl. NVLink): 30.1% of unexpected interruptions
HBM3 memory faults: 17.2% of unexpected interruptions
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs (NCSA Delta / DeltaAI)
1,056 GPUs (A100 + H100), 11.7M GPU-hours over 2.5 years · 2.5 years of production operation
H100 vs A100 memory errors: H100 shows ~3.2× LOWER per-GPU mean-time-between-(memory)-errors than A100. HBM3 error handling has not kept pace with capacity
H100 core hardware: H100 shows improved resilience vs A100 on critical (non-memory) hardware components
Capacity planning: ~5% GPU overprovisioning projected necessary at scale to absorb failures
What we deliberately do NOT publish yet: a Denpex-fleet MTBF table. Cross-tenant incident telemetry only began accumulating in July 2026, and an index computed from weeks of data would be noise sold as signal. When the corpus crosses 90 days and the cross-tenant disclosure ships, this section lights up with real numbers and full denominators.
The documented failure surface
Counts of DISTINCT documented failure signatures in the public Denpex Failure Encyclopedia (559 published classes). These measure how many documented, fixable ways each slice of the stack fails, a coverage/complexity map for procurement diligence, not a field failure rate. Every count links back to entries any engineer can read in the Failure Encyclopedia - this is auditable, not vibes. 61 of 559 signatures are hardware-class; 13 NVIDIA Xid codes are covered (13, 31, 43, 48, 61, 62, 63, 64, 69, 74, 79, 94, 119).
By GPU family
How many DISTINCT documented ways training fails on/around each GPU family, a complexity map for diligence, not a field failure rate.
Slice
Documented failure signatures
Critical-severity
Where they concentrate
H100 (SXM/PCIe)
10
3
Reliability (2) · Hardware (2) · Communication (2)
A100
4
1
Reliability (2) · Hardware (2)
H200
2
1
Reliability (1) · Communication (1)
B200 / GB200
2
1
Reliability (1) · Communication (1)
By interconnect
NVLink/NVSwitch vs InfiniBand vs RoCE vs EFA vs PCIe, where the fabric bites.
Slice
Documented failure signatures
Critical-severity
Where they concentrate
PCIe
22
8
Hardware (14) · Communication (2) · Memory (1)
NVLink / NVSwitch
18
6
Hardware (8) · Communication (4) · Fail-Slow (1)
InfiniBand
17
3
Communication (10) · Network (2) · Distributed Communication (1)
RoCE / Ethernet
7
2
Communication (4) · Network (2) · Hardware (1)
AWS EFA
3
1
Communication (3)
By provider stack
Failure signatures that reference each provider's stack (EFA, GKE, AKS, Slurm/on-prem...). Higher ≠ worse cloud, it can mean more usage and better-documented failure modes. Read with the studies above.
Slice
Documented failure signatures
Critical-severity
Where they concentrate
Kubernetes (any cloud)
10
0
Infrastructure (5) · Communication (2) · Reliability (1)
On-prem / HPC (Slurm etc.)
8
1
Infrastructure (4) · Distributed Training (1) · Environment (1)
AWS
7
1
Communication (3) · Infrastructure (2) · Network (1)
Azure
3
0
Infrastructure (3)
Google Cloud
2
0
Infrastructure (1) · Network (1)
CoreWeave
2
0
Reliability (1) · Infrastructure (1)
RunPod
2
0
Infrastructure (2)
Lambda
1
1
Infrastructure (1)
Take the dataset with you
Machine-readable JSON (free, no key): GET https://api.denpex.com/api/reliability-index. Cite it in your procurement memo; the sources above carry the field numbers.