Skip to content

Mixture-of-experts dispatch and combine collectives time out on large NVLink domains

Expert-parallel routing sends every token to the ranks holding its chosen experts and gathers the results back, an all-to-all exchange whose volume depends on what the router chose. On large NVLink domains this pair of collectives can stop completing, and the job stalls inside it with no rank reporting a fault.

Quick answer

Separate dispatch from combine first, then re-run at a smaller domain size. If it only stalls at full scale, the number of concurrent peer exchanges is the factor rather than the model or the routing.

Symptom
File ".../bench/ep_harness.py", line 248, in sample
Root cause
It frequently does not present as a timeout at all, and that is the hardest part of recognising it. When a collective is abandoned the communicator is torn down while its kernels are still resident on the device. The next call that inspects device state, commonly torch.
Recommended fix
Establish whether the stall is in dispatch or in combine, since they are separate exchanges with different volumes and narrowing to one halves the search.
How Denpex helps
Denpex matches Mixture-of-experts dispatch and combine collectives time out on large NVLink domains across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
Communication#mixture-of-experts#expert-parallel#all-to-all#dispatch-combine#nvl72#collective-timeout

What this failure is

A stall in the all-to-all exchanges that implement mixture-of-experts routing, in which a dispatch or combine receive never completes on a large NVLink domain and the job halts inside the collective without any rank raising a fault.

Live diagnosis, no signup

Is this what broke your run? Paste your log.

You're reading about Mixture-of-experts dispatch and combine collectives time out on large NVLink domains. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.

training_logs.txt
No log to hand? Try one:

3 free diagnoses/day

Why it happens (the mechanism)

Expert parallelism turns communication into a function of the data. The router decides which ranks exchange how much, so no two steps look alike and the pattern is neither uniform nor symmetric. All-to-all already puts every rank in contact with every other, and a rack-scale domain multiplies those simultaneous contacts, so assumptions that were safe inside one node meet a regime they were never exercised in.

What you'll observe

  • The stall is inside a collective, so no rank has an error of its own to report
  • It depends on the data, because routing decides the exchange volume, and so it appears irregularly
  • It emerges at large domain sizes and cannot be reproduced on a smaller topology
  • Both the routing and the transport look correct in isolation

Common symptoms and what they mean

SymptomWhy it happens
torch.AcceleratorError: CUDA error: unspecified launch failure raised from torch.cuda.synchronize()It frequently does not present as a timeout at all, and that is the hardest part of recognising it. When a collective is abandoned the communicator is torn down while its kernels are still resident on the device. The next call that inspects device state, commonly torch.cuda.synchronize at the end of a benchmark iteration, is the one that collects the wreckage, and it reports an unspecified launch failure. The frame in the traceback is therefore innocent, and the exchange that stalled has already been cleaned up by the time anything is printed.
The synchronize call being the frame that reports the fault, several operations after the exchange that actually stalledExpert routing makes communication data-dependent. Which ranks exchange how much is decided per step by the router, so the traffic pattern changes from step to step and is not the uniform, symmetric shape that collective implementations are usually tuned and tested against.
No collective named anywhere in the traceback, because the failing call is an ordinary device synchronisationAll-to-all is the most demanding of those patterns because every rank talks to every other simultaneously. Scaling the domain multiplies the number of concurrent peer exchanges rather than adding to it, so resource limits and ordering assumptions that hold within a node can fail across a rack.
A receive operation timing out during the dispatch or combine phase of expert routingThe receive side is where it surfaces because that is the side that waits. A dispatch that was not sent, or was sent to a peer that had already moved on, leaves a receiver blocked with nothing to report except that its data did not arrive.
The stall occurring inside an all-to-all exchange rather than an all-reduceIt frequently does not present as a timeout at all, and that is the hardest part of recognising it. When a collective is abandoned the communicator is torn down while its kernels are still resident on the device. The next call that inspects device state, commonly torch.cuda.synchronize at the end of a benchmark iteration, is the one that collects the wreckage, and it reports an unspecified launch failure. The frame in the traceback is therefore innocent, and the exchange that stalled has already been cleaned up by the time anything is printed.
Onset only at large NVLink domain sizes such as a full rack-scale systemExpert routing makes communication data-dependent. Which ranks exchange how much is decided per step by the router, so the traffic pattern changes from step to step and is not the uniform, symmetric shape that collective implementations are usually tuned and tested against.

Which systems are affected

  • Mixture-of-experts models using expert parallelism, where routing drives an all-to-all
  • Rack-scale NVLink domains, where the number of simultaneous peer exchanges is far larger than in a single node
  • Low-latency dispatch and combine kernels tuned for such domains
  • Any workload whose collective volume is decided by data rather than fixed by shape

How to confirm this is the problem

Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.

  • ✓Treat an unspecified launch failure at a synchronisation point as a downstream report, not a cause. Check whether an expert-parallel exchange preceded it in the same iteration before investigating kernels.
  • ✓Determine whether the timing-out operation is the dispatch or the combine exchange; they are separate collectives and only one of them will be at fault.
  • ✓Run the same model and data at a smaller domain size. Completing there and stalling at full scale indicates the number of concurrent peer exchanges is the factor.
  • ✓Substitute a general-purpose collective for the specialised low-latency path and re-run, which isolates the kernel from the routing.

Root cause

  • It frequently does not present as a timeout at all, and that is the hardest part of recognising it. When a collective is abandoned the communicator is torn down while its kernels are still resident on the device. The next call that inspects device state, commonly torch.cuda.synchronize at the end of a benchmark iteration, is the one that collects the wreckage, and it reports an unspecified launch failure. The frame in the traceback is therefore innocent, and the exchange that stalled has already been cleaned up by the time anything is printed.
  • Expert routing makes communication data-dependent. Which ranks exchange how much is decided per step by the router, so the traffic pattern changes from step to step and is not the uniform, symmetric shape that collective implementations are usually tuned and tested against.
  • All-to-all is the most demanding of those patterns because every rank talks to every other simultaneously. Scaling the domain multiplies the number of concurrent peer exchanges rather than adding to it, so resource limits and ordering assumptions that hold within a node can fail across a rack.
  • The receive side is where it surfaces because that is the side that waits. A dispatch that was not sent, or was sent to a peer that had already moved on, leaves a receiver blocked with nothing to report except that its data did not arrive.

The fix and how to prevent it

Searchable error signature

search key
  File ".../bench/ep_harness.py", line 248, in sample
    torch.cuda.synchronize()
torch.AcceleratorError: CUDA error: unspecified launch failure
dispatch/combine receives time out
[rank17] Watchdog caught collective operation timeout: WorkNCCL(OpType=ALLTOALL) ran for 1800000 milliseconds before timing out

Use this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.

The fix and the prevention pattern

The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.

Sign up free. Unlock the full analysis

No credit card. Daily allowance follows verified trust tier. Instant access.

Why the recommended fix works

Narrowing to dispatch or combine, and to a domain size, replaces a whole-system symptom with a specific exchange at a specific scale, the only form in which it can be investigated. Falling back to a general-purpose collective separates the specialised kernel from the routing, and capping expert load removes the extreme imbalances that make the exchange hardest.

Best practices by model family

Model / StackRecommendationNotes
Stall inside expert routingSeparate dispatch from combine firstThey are distinct exchanges; only one will be at fault.
Suspected scale dependenceRe-run at a smaller domain sizeCompleting smaller and stalling larger implicates concurrent peer count.
Low-latency kernels in useFall back to a general-purpose collectiveIsolates the specialised path from the routing logic.
Irregular, data-dependent onsetLog the routing distribution per stepWithout it an unusual step cannot be correlated with the stall.

With the fix vs without the fix

DimensionWith the fixWithout the fix
What decides the trafficThe router, per step, from the dataAssumed fixed by the model shape
Effect of domain sizeMultiplies simultaneous peer exchangesAssumed to add capacity linearly
Which side reportsThe receiver, which is waitingExpected from whichever side failed

Diagnostic note

“The data dependence is what makes this so hard to pin down, and it is worth saying plainly: the same model, the same code and the same hardware will complete thousands of steps and stall on one, because the router produced an unusual distribution on that step. Anyone treating it as a flaky interconnect will chase it indefinitely. Capture the routing distribution per step, or the stall stays unattributable.”

Visual fingerprint

Why the pattern changes every step
  step N     router sends       rank0 -> {2,5,9}   rank1 -> {2}      rank2 -> {0,1,...}
  step N+1   router sends       rank0 -> {3}       rank1 -> {3,7,8}  rank2 -> {5}

  all-to-all: every rank in contact with every other, volumes set by routing
  larger domain -> more simultaneous exchanges, not merely more capacity
Routing decides which ranks exchange how much on every step, so the communication pattern is different each time and is neither uniform nor symmetric. Growing the domain increases the number of simultaneous peer exchanges rather than simply providing more bandwidth.

Diagnose this failure in VS Code

Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.

Install the free VS Code extension

Frequently asked questions

Questions engineers and on-call staff commonly ask about this failure.

Why does it only happen at full scale?
All-to-all puts every rank in contact with every other, so growing the domain multiplies simultaneous exchanges. Assumptions that hold inside one node meet a regime they were not exercised in.
The same job ran fine a thousand times.
Routing is data-dependent, so the exchange pattern differs every step. An unusual distribution on one step is enough, which is why it looks like flaky hardware.
No rank reported an error.
The stall is inside a collective. Receivers are waiting and have nothing to report except that data did not arrive; nothing raised a fault.
Could this be a hardware fault instead?
Yes, and it looks identical. Check the host kernel logs for an accelerator fault at the moment progress stopped before pursuing the routing path.

Don't just read the fix, diagnose your run

The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.