Skip to content
Evidence and evaluation transparency

What Denpex can prove today.

Diagnostic claims are easy to overstate. This page separates internal development measurements from independent evidence and real customer outcomes, including the gaps that still prevent us from publishing an accuracy percentage.

Download the verified-fix postmortem template (PDF, 6 pages)

Current evidence status

A

Development benchmark

Internal evidence

Labels were hidden during scoring, but the incidents were previously visible during engine development. The result is useful for regression testing, not customer accuracy.

B

Sealed holdout

Not yet published

Zero independently acquired production holdout results are currently published. Until that changes, Denpex does not present an unseen-incident accuracy rate.

C

Customer outcomes

No public cohort

Denpex records authenticated fix feedback and recovery verification for each tenant's operational loop. We do not currently publish aggregate customer outcome statistics.

RMA evidence packet

This seeded packet shows the evidence boundary before hardware escalation. One uncorrectable ECC event earns a field-diagnostic verdict, not an automatic replacement request.

Verdict tier

FIELD_DIAG_FIRST

Uncorrectable memory error observed. Capture a bug report before an approved reset, check row-remap status, then run diagnostics on an isolated GPU. The vendor determines RMA eligibility.

Single observed Xid event. Recurrence is not established.

Device identity

vendor
NVIDIA
hostname
node-05
gpu Uuid
GPU-9f8e7d6c-1234-4abc-9def-0123456789ab
gpu Serial
1560921003217
pci Bus Id
0000:8a:00
driver Version
550.90.07
vbios Version
96.00.99.00.01

Missing evidence before submission

  • Missing: System / chassis serial (service tag)
  • Missing: nvidia-bug-report archive (NVIDIA delays RMAs without it)
  • Missing: DCGM diagnostic result
  • Missing: NVIDIA Field Diagnostic result
  • Missing: BMC system event log (server-vendor cases)

Verification controls

  1. nvidia-smi -q -d ROW_REMAPPER # check "Remapping Failure Occurred" and pending remaps
  2. sudo nvidia-bug-report.sh # capture evidence before any reset or power cycle
  3. Manual approval required before draining node-05; isolate workloads before diagnostics.
  4. Manual approval required before resetting GPU 00000000; confirm that no workload owns it.
  5. Manual approval required before dcgmi diag -r 3; run it only on an isolated idle GPU.
  6. Manual approval required before dcgmi diag -r 4; use extended stress only on an isolated idle GPU.
  7. Request NVIDIA Field Diagnostic from the vendor before deciding on RMA eligibility.

Support-ready text

Vendor acceptance is not claimed. This packet organizes support evidence, but the vendor makes the warranty and replacement decision.

GPU HARDWARE REVIEW DRAFT (denpex.rma.v2)
-----------------------------------------------------------
Verdict: FIELD_DIAG_FIRST
Reason: Uncorrectable memory error observed. Capture a bug report before an approved reset, check row-remap status, then run diagnostics on an isolated GPU. The vendor determines RMA eligibility.
Vendor: NVIDIA
Host: node-05   System serial: Not found, run dmidecode
GPU UUID: GPU-9f8e7d6c-1234-4abc-9def-0123456789ab
GPU serial: 1560921003217
PCI bus: 0000:8a:00   Driver: 550.90.07   VBIOS: 96.00.99.00.01
Time of failure: Jun 10 08:14:19
Fault signature: GPU_XID_48   Xid codes: 48
Evidence checklist:
  [x] GPU UUID
  [x] GPU serial number
  [ ] System / chassis serial (service tag), collect via: sudo dmidecode -s system-serial-number
  [x] Driver + VBIOS versions
  [x] Xid excerpt with timestamps
  [x] ECC / row-remap counters
  [ ] nvidia-bug-report archive (NVIDIA delays RMAs without it), collect via: sudo nvidia-bug-report.sh   # run on the failing node, attach nvidia-bug-report.log.gz
  [ ] DCGM diagnostic result, collect via: dcgmi diag -r 3   # run on an isolated idle GPU
  [ ] NVIDIA Field Diagnostic result, collect via: Request the Field Diagnostic procedure and tool from NVIDIA or your hardware vendor
  [ ] BMC system event log (server-vendor cases), collect via: ipmitool sel elist | tail -50
Xid events:
  Jun 10 08:14:19 node-05 kernel: NVRM: Xid (PCI:0000:8a:00): 48, pid=4012, name=python, Double Bit ECC Error
ECC / remap evidence:
  Jun 10 08:14:19 node-05 kernel: NVRM: Xid (PCI:0000:8a:00): 48, pid=4012, name=python, Double Bit ECC Error
  [rank42] ECC uncorrectable error detected on GPU 00000000:8A:00.0
This is an evidence draft, not an RMA approval. Run the verification flow and ask the vendor to determine eligibility.
Generated by Denpex ML Incident Intelligence.

Evidence provenance

Seeded public demonstration derived from committed RMA regression cases. It contains no customer or vendor case data.

Tenant memory and verified recovery

This seeded replay shows how a team's authenticated prior fix is replayed alongside the engine recommendation, attributed and advisory rather than overwriting it, while recovery remains open until the required server-clock observation window finishes.

Prior confirmed fix replay

Confirmed 3 times

Your team's confirmed fix for NCCL_TIMEOUT (last confirmed 2026-08-20 by seeded-operator@example.invalid, confirmed 3 times):

Pin NCCL_IB_HCA to mlx5_0 and rerun the same collective control.

Engine recommendation

Engine recommendation: inspect the first transport error and compare a known-good rail.

Fix source
team-kb (advisory)
Last confirmed
2026-08-20

Verification report

Awaiting observation

Recovery remains open. A passing no-recurrence report is still pending until its full 30-minute window elapses.

  • The GPU enumerates again through NVMLpassed

    Required. Observation window: 0 minutes.

  • The same incident did not recurpending

    Required. Observation window: 30 minutes.

Tenant isolation

This confirmed fix belongs to seeded Team Alpha only. It is not public knowledge and is not available to another tenant.

Evidence provenance

Seeded public demonstration of production team-memory and verification contracts. It contains no customer data.

First-rank localization with telemetry

In this seeded fixture, rank 3 reports the first primary fault. Ranks 0, 1, 2 report later collective timeouts and are collateral victims. The telemetry overlay records one volatile uncorrectable ECC event on the initiating GPU.

One GPU fault, three collateral rank timeouts
TimeRank and nodeCausal roleEvidence
12:00:00.120 UTCRank 3, node-bInitiator[rank3] CUDA error: uncorrectable ECC error encountered
12:00:02.100 UTCRank 0, node-aCollateral victim[rank0] Watchdog caught collective operation timeout
12:00:02.220 UTCRank 1, node-aCollateral victim[rank1] NCCL timeout waiting for all-reduce
12:00:02.310 UTCRank 2, node-aCollateral victim[rank2] collective operation timeout

Evidence provenance

Seeded public demonstration derived from committed deterministic regression cases. It is not a customer incident or an accuracy sample.

Uncertainty boundary

This fixture demonstrates one covered causal pattern. It does not establish correctness for incomplete incidents, other failure families, or production populations.

Development evidence, with its limitation attached

What was measured

A reproducible comparison scores the Denpex causal pipeline and three specified algorithmic heuristics on the same development corpus. It is useful for detecting regressions and checking whether causal ranking improves on simple signature selection.

What it does not establish

It is not a sealed holdout, a production accuracy estimate, a comparison with ML infrastructure engineers, or a comparison with general AI assistants. We therefore do not promote its class-match percentage as customer proof.

Reproduction command: npx tsx worker/eval/run-blinded-comparison.ts

The bar for publishable accuracy evidence

Denpex will not call an evaluation independent merely because labels were hidden during one run. A publishable sealed holdout must satisfy every requirement below.

  • Incidents the diagnostic engine and its authors have never used for tuning
  • Maintainer-confirmed or operator-confirmed root causes and resolutions
  • Semantic grading of the diagnosis, evidence, commands, verification, and abstention
  • Coverage across failure families, frameworks, schedulers, accelerators, and incomplete evidence
  • A frozen evaluation protocol and published sample counts, limitations, and retired cases

Customer-observed outcomes

No consented customer outcome sample is available for public reporting yet. A future public result must separate human-confirmed outcomes from automated checks, suppress small cohorts, prevent duplicate submissions, and honor deletion and consent boundaries.

Until those controls and a sufficient independent cohort are production-verified, this page will remain an explicit empty state instead of displaying zeroes as if they were evidence.

Inspect the product directly

Run the public matcher or review the architecture and safety boundaries behind the diagnosis pipeline.