Tenant job or fleet fault?
Separate application and image failures from infrastructure events before draining capacity or moving the customer.
For GPU clouds and neocloud operators
GPU cloud reliability is not just uptime. The hard problem is deciding whether a failed tenant job belongs to the workload, image, fabric, host, slot or GPU before another customer lands on the same node.
Denpex combines first-failure log analysis with GPU telemetry and incident history. Operators get an evidence-backed diagnosis, a verification step, and a hardware escalation path that distinguishes an RMA candidate from a healthy card on a bad slot or node.
These incidents look adjacent in the final alert, but they require different owners and different proof before recovery.
Separate application and image failures from infrastructure events before draining capacity or moving the customer.
Correlate Xid, ECC, NVLink, reset and device-loss evidence by GPU UUID, node and time window.
Find the earliest distinct rank or node instead of opening one incident for every rank that later timed out.
Require vendor-grade device identity, recurrence and verification evidence before a replacement case is assembled.
Wrap customer-visible jobs or ingest scheduler and telemetry events without requiring a kernel module.
Order cross-rank failures, identify the initiating event, and map it to the affected node and GPU identity.
Run a controlled check that separates workload, node, slot and card ownership before returning capacity.
Retain tenant-scoped incident outcomes so repeated nodes and failure families become measurable fleet reliability signals.
Use your incidents and timestamps. These are measurement definitions, not Denpex customer-outcome claims.
| Metric | Measurement |
|---|---|
| Time to causal owner | Minutes from first failure to workload, fabric, host, slot or GPU assignment, compared with the current support process. |
| False RMA avoidance | Cards withheld from replacement after the controlled test showed a node, slot, power or software cause. |
| GPU hours recovered | Idle or failed GPU time avoided after verified recovery, using actual allocation timestamps rather than a modeled estimate. |
| Support deflection | Incidents resolved by operators with complete evidence before a tenant or vendor escalation was opened. |
Classify common Xid codes and see which ones can support an RMA investigation.
Review sourced field evidence and failure surfaces by accelerator and interconnect.
Separate transport and fabric failures from the ranks that only reported the cascade.
Export tenant-scoped failures, MTTR, Xid events and measured GPU-hour loss.
Denpex can organize the evidence and require the discriminating control, such as whether the failure follows the card to a known-good slot. A single Xid 79 does not prove card ownership, so the workflow preserves power, PCIe, BMC and recurrence evidence before an RMA verdict.
No. Diagnosis uses logs, failure signatures, process metadata and optional GPU telemetry. Client-side masking is the default, strict mode sends only anonymized failure signatures, and local mode can keep the diagnosis path and incident record inside the operator environment.
Replay a fixed set of resolved incidents plus live shadow traffic, freeze the expected causal owner, and measure diagnosis time, evidence completeness, false escalations and verified recovery. The evaluation page provides a protocol that does not require trusting a vendor accuracy percentage.
The standard public agent provides monitoring, diagnosis, dry runs and human-run verification steps. Hands-off remote execution is limited to qualified controlled pilots with scoped credentials, authorization, readback verification and rollback.
Freeze the expected owner and outcome, replay the evidence, then compare operator time and decision quality with your current process.
Build the evaluation plan