Skip to content

Evaluation and pilot

Prove GPU failure diagnosis on incidents your team already resolved

A credible evaluation does not ask you to trust a benchmark or a polished demo. Freeze the answer first, replay the evidence operators really had, and measure decision quality, time and recovery against your current process.

Four evaluation phases

  1. 1

    Freeze the incident set

    Select resolved incidents that Denpex has not been tuned on. Record the operator-confirmed initiating owner, correction and recovery evidence before any Denpex result is opened.

  2. 2

    Replay historical evidence

    Submit the evidence operators actually had at the time, not the clean postmortem. Score the diagnosis, missing-evidence request, action and verification separately.

  3. 3

    Run live in shadow mode

    Collect diagnoses alongside the current process without granting action authority. Compare timestamps, routing, escalations and operator decisions on the same incidents.

  4. 4

    Verify the outcome

    Require the initiating signature to disappear under a production-shaped control. A command that exited zero counts as attempted recovery until readback and the observation window pass.

Score each dimension separately

A diagnosis can name the right class while recommending the wrong action. One blended accuracy number hides that difference.

Evaluation dimensions and their scoring questions.
DimensionQuestion
Causal-owner matchDid the result assign the incident to the same workload, config, runtime, fabric, host, slot or GPU owner confirmed by the final investigation?
Evidence completenessDid it cite the first failure, rank or node identity, relevant configuration and the evidence needed to rule out adjacent causes?
Action safetyWas the first action specific, reversible where possible, scoped to the affected resource and free of unsupported destructive steps?
Abstention qualityWhen the evidence was incomplete, did the result preserve competing hypotheses and request the cheapest discriminating artifact?
Time to decisionHow many engineer minutes elapsed from incident availability to a verified owner and next action under each process?
Verified recoveryDid the same production-shaped control pass, and did the initiating signature remain absent through the agreed observation window?

Outcome metrics a buyer can defend

These definitions produce customer-owned measurements. Denpex does not present them as public customer results until the customer consents and the cohort is supportable.

Incident time saved

Baseline operator minutes minus evaluation operator minutes for the same incident stage and evidence.

False RMA avoidance

Suspected boards withheld after a controlled test proved software, slot, node, power, thermal or fabric ownership.

GPU hours recovered

Measured allocation time restored or protected after verified correction, based on scheduler and agent timestamps.

Support tickets deflected

Incidents resolved to the agreed standard without an additional platform, vendor or hardware escalation.

Pilot guardrails

  • Keep the scored incident cohort separate from implementation and tuning examples.
  • Begin with read-only diagnosis and shadow routing before any action authority.
  • Use one agreed reviewer rubric and preserve disagreements instead of forcing consensus.
  • Separate measured runtime and labor from modeled opportunity or hardware cost.
  • Define rollback, stop conditions, deletion and access boundaries before live collection.

Frequently asked questions

Does Denpex publish a production accuracy percentage?

No. The public proof page distinguishes internal development measurements from a sealed independent holdout and customer outcomes. Until a qualifying holdout exists, the evaluation should use your frozen incidents and operator-confirmed outcomes instead of a vendor accuracy percentage.

Can Denpex tune on our evaluation incidents?

Not if the result is meant to test generalization. Freeze the cohort and expected outcomes before replay, keep it separate from any implementation or tuning set, and record any incident that Denpex or its authors previously saw as ineligible for the sealed score.

Do we need to send raw production logs?

No. Teams can begin with consented and redacted historical evidence, use client-side masking or strict signature-only mode, or evaluate the local deterministic path inside the environment. The selected evidence level should be recorded because it affects what can be concluded.

What turns an evaluation into a production rollout?

The parties agree on passed technical and operational criteria, deployment and data-flow approval, named owners, alert routing, rollback, support boundaries and a measured expansion cohort. Remote execution remains separately gated and is not implied by a diagnosis pilot.

Bring the incidents that cost your team the most time

We will scope the evidence path, deployment boundary, reviewer rubric and outcome metrics before any result is scored.

Discuss the evaluation