Skip to content
Evidence and evaluation transparency

What Denpex can prove today.

Diagnostic claims are easy to overstate. This page separates internal development measurements from independent evidence and real customer outcomes, including the gaps that still prevent us from publishing an accuracy percentage.

Current evidence status

A

Development benchmark

Internal evidence

Labels were hidden during scoring, but the incidents were previously visible during engine development. The result is useful for regression testing, not customer accuracy.

B

Sealed holdout

Not yet published

Zero independently acquired production holdout results are currently published. Until that changes, Denpex does not present an unseen-incident accuracy rate.

C

Customer outcomes

No public cohort

Denpex records authenticated fix feedback and recovery verification for each tenant's operational loop. We do not currently publish aggregate customer outcome statistics.

Development evidence, with its limitation attached

What was measured

A reproducible comparison scores the Denpex causal pipeline and three specified algorithmic heuristics on the same development corpus. It is useful for detecting regressions and checking whether causal ranking improves on simple signature selection.

What it does not establish

It is not a sealed holdout, a production accuracy estimate, a comparison with ML infrastructure engineers, or a comparison with general AI assistants. We therefore do not promote its class-match percentage as customer proof.

Reproduction command: npx tsx worker/eval/run-blinded-comparison.ts

The bar for publishable accuracy evidence

Denpex will not call an evaluation independent merely because labels were hidden during one run. A publishable sealed holdout must satisfy every requirement below.

  • Incidents the diagnostic engine and its authors have never used for tuning
  • Maintainer-confirmed or operator-confirmed root causes and resolutions
  • Semantic grading of the diagnosis, evidence, commands, verification, and abstention
  • Coverage across failure families, frameworks, schedulers, accelerators, and incomplete evidence
  • A frozen evaluation protocol and published sample counts, limitations, and retired cases

Customer-observed outcomes

No consented customer outcome sample is available for public reporting yet. A future public result must separate human-confirmed outcomes from automated checks, suppress small cohorts, prevent duplicate submissions, and honor deletion and consent boundaries.

Until those controls and a sufficient independent cohort are production-verified, this page will remain an explicit empty state instead of displaying zeroes as if they were evidence.

Inspect the product directly

Run the public matcher or review the architecture and safety boundaries behind the diagnosis pipeline.