Development benchmark
Internal evidence
Labels were hidden during scoring, but the incidents were previously visible during engine development. The result is useful for regression testing, not customer accuracy.
Diagnostic claims are easy to overstate. This page separates internal development measurements from independent evidence and real customer outcomes, including the gaps that still prevent us from publishing an accuracy percentage.
Internal evidence
Labels were hidden during scoring, but the incidents were previously visible during engine development. The result is useful for regression testing, not customer accuracy.
Not yet published
Zero independently acquired production holdout results are currently published. Until that changes, Denpex does not present an unseen-incident accuracy rate.
No public cohort
Denpex records authenticated fix feedback and recovery verification for each tenant's operational loop. We do not currently publish aggregate customer outcome statistics.
A reproducible comparison scores the Denpex causal pipeline and three specified algorithmic heuristics on the same development corpus. It is useful for detecting regressions and checking whether causal ranking improves on simple signature selection.
It is not a sealed holdout, a production accuracy estimate, a comparison with ML infrastructure engineers, or a comparison with general AI assistants. We therefore do not promote its class-match percentage as customer proof.
Reproduction command: npx tsx worker/eval/run-blinded-comparison.ts
Denpex will not call an evaluation independent merely because labels were hidden during one run. A publishable sealed holdout must satisfy every requirement below.
No consented customer outcome sample is available for public reporting yet. A future public result must separate human-confirmed outcomes from automated checks, suppress small cohorts, prevent duplicate submissions, and honor deletion and consent boundaries.
Until those controls and a sufficient independent cohort are production-verified, this page will remain an explicit empty state instead of displaying zeroes as if they were evidence.
Run the public matcher or review the architecture and safety boundaries behind the diagnosis pipeline.