Incident time saved
Baseline operator minutes minus evaluation operator minutes for the same incident stage and evidence.
Evaluation and pilot
A credible evaluation does not ask you to trust a benchmark or a polished demo. Freeze the answer first, replay the evidence operators really had, and measure decision quality, time and recovery against your current process.
Select resolved incidents that Denpex has not been tuned on. Record the operator-confirmed initiating owner, correction and recovery evidence before any Denpex result is opened.
Submit the evidence operators actually had at the time, not the clean postmortem. Score the diagnosis, missing-evidence request, action and verification separately.
Collect diagnoses alongside the current process without granting action authority. Compare timestamps, routing, escalations and operator decisions on the same incidents.
Require the initiating signature to disappear under a production-shaped control. A command that exited zero counts as attempted recovery until readback and the observation window pass.
A diagnosis can name the right class while recommending the wrong action. One blended accuracy number hides that difference.
| Dimension | Question |
|---|---|
| Causal-owner match | Did the result assign the incident to the same workload, config, runtime, fabric, host, slot or GPU owner confirmed by the final investigation? |
| Evidence completeness | Did it cite the first failure, rank or node identity, relevant configuration and the evidence needed to rule out adjacent causes? |
| Action safety | Was the first action specific, reversible where possible, scoped to the affected resource and free of unsupported destructive steps? |
| Abstention quality | When the evidence was incomplete, did the result preserve competing hypotheses and request the cheapest discriminating artifact? |
| Time to decision | How many engineer minutes elapsed from incident availability to a verified owner and next action under each process? |
| Verified recovery | Did the same production-shaped control pass, and did the initiating signature remain absent through the agreed observation window? |
These definitions produce customer-owned measurements. Denpex does not present them as public customer results until the customer consents and the cohort is supportable.
Baseline operator minutes minus evaluation operator minutes for the same incident stage and evidence.
Suspected boards withheld after a controlled test proved software, slot, node, power, thermal or fabric ownership.
Measured allocation time restored or protected after verified correction, based on scheduler and agent timestamps.
Incidents resolved to the agreed standard without an additional platform, vendor or hardware escalation.
No. The public proof page distinguishes internal development measurements from a sealed independent holdout and customer outcomes. Until a qualifying holdout exists, the evaluation should use your frozen incidents and operator-confirmed outcomes instead of a vendor accuracy percentage.
Not if the result is meant to test generalization. Freeze the cohort and expected outcomes before replay, keep it separate from any implementation or tuning set, and record any incident that Denpex or its authors previously saw as ineligible for the sealed score.
No. Teams can begin with consented and redacted historical evidence, use client-side masking or strict signature-only mode, or evaluate the local deterministic path inside the environment. The selected evidence level should be recorded because it affects what can be concluded.
The parties agree on passed technical and operational criteria, deployment and data-flow approval, named owners, alert routing, rollback, support boundaries and a measured expansion cohort. Remote execution remains separately gated and is not implied by a diagnosis pilot.
We will scope the evidence path, deployment boundary, reviewer rubric and outcome metrics before any result is scored.
Discuss the evaluation