Skip to content

For internal AI platforms

A reliability layer for ML platform and AI infrastructure teams

ML platform teams inherit failures from every layer without owning all of them. Researchers see a dead training job, while the platform team must decide whether the code, framework, container, scheduler, network, node or accelerator should receive the ticket.

Denpex creates a common incident language across those layers. It collects or accepts the available evidence, identifies the earliest causal event, supplies a verification step, and records the outcome so the same failure does not start from zero next time.

Who this is for

  • ML infrastructure engineers
  • AI platform engineers
  • MLOps engineers
  • Platform SREs
  • Distributed systems engineers
  • Directors of AI infrastructure

The failure surface

These incidents look adjacent in the final alert, but they require different owners and different proof before recovery.

Framework diversity

Normalize failures across PyTorch DDP and FSDP, DeepSpeed, vLLM, JAX and supporting CUDA or NCCL layers.

Scheduler and cloud boundaries

Keep Slurm, Kubernetes, Ray, container and provider evidence attached to the same incident.

Researcher support load

Give users an answer-first runbook and a clear escalation boundary instead of a generic request to send more logs.

Reliability measurement

Track verified recovery, repeat failures, GPU-hour loss and node-level recurrence from tenant-owned data.

The operating loop

  1. 1

    Standardize intake

    Use the agent, API, console, editor or saved-log path to collect the same core incident fields.

  2. 2

    Route by causal owner

    Send application, platform, fabric and hardware incidents to the team that can actually change the failing condition.

  3. 3

    Require a verification control

    Treat a command that exits successfully as attempted recovery until the expected observations pass.

  4. 4

    Review recurring cost

    Use incident history and measured runtime to prioritize repeated nodes, configurations and support topics.

What Denpex contributes

  • One diagnosis pipeline across console, API, agent, editor and local paths
  • Deterministic known-signature matching plus a deeper path for novel or ambiguous evidence
  • Alert, metrics and scheduler integration points for existing platform workflows
  • Tenant-scoped incident and verification history rather than shared customer data

Measure this in an evaluation

Use your incidents and timestamps. These are measurement definitions, not Denpex customer-outcome claims.

Recommended evaluation metrics and how to measure them.
MetricMeasurement
Support time per incidentEngineer minutes from ticket open to verified owner and next action, sampled by framework and failure family.
Ticket deflectionResearcher incidents resolved through the runbook without a platform engineer joining the investigation.
Repeat failure rateIncidents with an initiating signature and configuration already seen after the prior correction.
GPU hours lostMeasured failed allocation time from agent and scheduler events, kept separate from modeled opportunity cost.

Frequently asked questions

Does Denpex replace our observability platform?

No. Metrics and logs show system state, while Denpex focuses on causal failure diagnosis, evidence collection and verification. It can export incident metrics and attach existing observability evidence without requiring the platform team to replace Grafana, Datadog or its scheduler.

How does Denpex handle failures it has never seen?

Known signatures use the deterministic path. Novel or ambiguous cases can use a deeper evidence-permitted path, retain competing hypotheses, or request a discriminating artifact. They should not be silently promoted into deterministic coverage without review and regression tests.

Can different research teams keep their incidents isolated?

Customer history and metrics are tenant-scoped. Client-side masking, strict signature-only collection and local diagnosis provide additional data-minimization choices. Security and procurement should verify the selected deployment against the published trust materials.

Who normally owns a Denpex rollout?

The operational owner is usually ML infrastructure, AI platform, SRE or GPU fleet engineering. Security reviews data flow and deployment, while a director of AI infrastructure or engineering buyer defines the pilot outcome and expansion boundary.

Prove it on incidents your team already resolved

Freeze the expected owner and outcome, replay the evidence, then compare operator time and decision quality with your current process.

Build the evaluation plan