Skip to content

For GPU clouds and neocloud operators

GPU cloud reliability without blind node replacement

GPU cloud reliability is not just uptime. The hard problem is deciding whether a failed tenant job belongs to the workload, image, fabric, host, slot or GPU before another customer lands on the same node.

Denpex combines first-failure log analysis with GPU telemetry and incident history. Operators get an evidence-backed diagnosis, a verification step, and a hardware escalation path that distinguishes an RMA candidate from a healthy card on a bad slot or node.

Who this is for

  • GPU cloud reliability engineers
  • Fleet SREs
  • Bare-metal operations teams
  • Hardware qualification engineers
  • Support escalation leads
  • Capacity and FinOps teams

The failure surface

These incidents look adjacent in the final alert, but they require different owners and different proof before recovery.

Tenant job or fleet fault?

Separate application and image failures from infrastructure events before draining capacity or moving the customer.

Intermittent GPU and PCIe faults

Correlate Xid, ECC, NVLink, reset and device-loss evidence by GPU UUID, node and time window.

Fabric-wide cascades

Find the earliest distinct rank or node instead of opening one incident for every rank that later timed out.

RMA evidence quality

Require vendor-grade device identity, recurrence and verification evidence before a replacement case is assembled.

The operating loop

  1. 1

    Observe

    Wrap customer-visible jobs or ingest scheduler and telemetry events without requiring a kernel module.

  2. 2

    Localize

    Order cross-rank failures, identify the initiating event, and map it to the affected node and GPU identity.

  3. 3

    Verify

    Run a controlled check that separates workload, node, slot and card ownership before returning capacity.

  4. 4

    Learn

    Retain tenant-scoped incident outcomes so repeated nodes and failure families become measurable fleet reliability signals.

What Denpex contributes

  • Deterministic coverage for Xid, CUDA, NCCL, scheduler, container and framework failures
  • Per-node reliability, incident history, measured GPU-hour loss and RMA evidence workflows
  • Client-side log masking, strict signature-only mode, and in-VPC or local diagnosis paths
  • Prometheus metrics for existing fleet operations and support dashboards

Measure this in an evaluation

Use your incidents and timestamps. These are measurement definitions, not Denpex customer-outcome claims.

Recommended evaluation metrics and how to measure them.
MetricMeasurement
Time to causal ownerMinutes from first failure to workload, fabric, host, slot or GPU assignment, compared with the current support process.
False RMA avoidanceCards withheld from replacement after the controlled test showed a node, slot, power or software cause.
GPU hours recoveredIdle or failed GPU time avoided after verified recovery, using actual allocation timestamps rather than a modeled estimate.
Support deflectionIncidents resolved by operators with complete evidence before a tenant or vendor escalation was opened.

Frequently asked questions

Can Denpex tell whether a GPU or the server slot is at fault?

Denpex can organize the evidence and require the discriminating control, such as whether the failure follows the card to a known-good slot. A single Xid 79 does not prove card ownership, so the workflow preserves power, PCIe, BMC and recurrence evidence before an RMA verdict.

Does Denpex require access to tenant model weights or training data?

No. Diagnosis uses logs, failure signatures, process metadata and optional GPU telemetry. Client-side masking is the default, strict mode sends only anonymized failure signatures, and local mode can keep the diagnosis path and incident record inside the operator environment.

How should a GPU cloud evaluate Denpex?

Replay a fixed set of resolved incidents plus live shadow traffic, freeze the expected causal owner, and measure diagnosis time, evidence completeness, false escalations and verified recovery. The evaluation page provides a protocol that does not require trusting a vendor accuracy percentage.

Can Denpex automatically replace or reset customer GPUs?

The standard public agent provides monitoring, diagnosis, dry runs and human-run verification steps. Hands-off remote execution is limited to qualified controlled pilots with scoped credentials, authorization, readback verification and rollback.

Prove it on incidents your team already resolved

Freeze the expected owner and outcome, replay the evidence, then compare operator time and decision quality with your current process.

Build the evaluation plan