vLLM runs out of GPU memory during serving after the prefill token budget is raised
Raising the number of tokens vLLM may batch per step improves time to first token and throughput, and it also raises the peak activation memory a prefill step needs. When that peak collides with the memory already reserved for the KV cache, the failure arrives during serving rather than at startup, after the deployment has looked healthy.
Drop --max-num-batched-tokens to 2048 to 4096 first, then --max-num-seqs, and only then touch gpu_memory_utilization. A runtime failure after a throughput change points at step size, not at cache size.
- Symptom
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 3.83 GiB. GPU 0 has a total capacity of 79.15 GiB of which 2.11 GiB is free.- Root cause
- Two different budgets draw on the same device memory. gpu_memory_utilization reserves a pool that the KV cache is carved from, while the batched token limit governs how much work a single step may do and therefore how large the activation tensors for that step become. Raising the second does not consult the first.
- Recommended fix
- Lower --max-num-batched-tokens toward 2048 to 4096 as the first move when the failure appears at runtime rather than at startup. This is the setting that grew the step.
- How Denpex helps
- Denpex matches vLLM runs out of GPU memory during serving after the prefill token budget is raised across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
What this failure is
A serving-time out-of-memory failure caused by the per-step prefill token budget rather than by the KV cache reservation, appearing only once arriving prompts are long enough to build a maximum-sized batch.
Is this what broke your run? Paste your log.
You're reading about vLLM runs out of GPU memory during serving after the prefill token budget is raised. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.
Why it happens (the mechanism)
The batched token budget decides how much prefill work goes into one step, and activation memory for that step scales with it. That memory is not part of the pool the KV cache was sized against, so raising the budget quietly increases peak usage above what startup profiling measured. Profiling runs before any real prompt has arrived, so the configuration that fails under long prompts is one that was never exercised at its own limit.
What you'll observe
- The server starts cleanly and fails only once real traffic arrives
- The failure correlates with long prompts rather than with the number of concurrent users
- A configuration tuned for throughput on one model fails on another of similar size
- Reducing the served context length does not help, because the prompts were already within it
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| torch.OutOfMemoryError: CUDA out of memory raised during a forward pass while serving rather than during startup | Two different budgets draw on the same device memory. gpu_memory_utilization reserves a pool that the KV cache is carved from, while the batched token limit governs how much work a single step may do and therefore how large the activation tensors for that step become. Raising the second does not consult the first. |
| Failures beginning immediately after --max-num-batched-tokens or --enable-chunked-prefill was changed | The token budget only binds when a step is large enough to reach it, which requires long prompts or a heavy prefill mix. A deployment can therefore pass startup profiling, serve ordinary traffic for hours, and fail the first time the arriving prompts are long enough to build a full-sized batch. |
| A log line advising that if out-of-memory occurs during cudagraph capture, consider decreasing gpu_memory_utilization or switching to eager mode | The reason the served context length is not the lever is that the prompts were always inside it. What changed is how many prefill tokens the scheduler is willing to put into one step, which is a throughput setting rather than a capacity one. |
Which systems are affected
- vLLM V1, where chunked prefill is enabled by default whenever possible
- Deployments tuned toward a batched token budget above 8192 for throughput
- Mixed traffic combining long prompts with many short decodes in the same step
- Configurations where max_num_seqs multiplied by max_model_len already consumes most of the cache budget
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Check whether the failure began with a change to --max-num-batched-tokens or --enable-chunked-prefill rather than with a change in traffic volume.
- ✓Correlate the failures against prompt length rather than request rate. A token budget only binds once a step is large enough to reach it.
- ✓Restart with the budget set to 2048 and replay the same traffic. Success at the lower budget with everything else unchanged confirms the step size rather than the cache was the cause.
Root cause
- Two different budgets draw on the same device memory. gpu_memory_utilization reserves a pool that the KV cache is carved from, while the batched token limit governs how much work a single step may do and therefore how large the activation tensors for that step become. Raising the second does not consult the first.
- The token budget only binds when a step is large enough to reach it, which requires long prompts or a heavy prefill mix. A deployment can therefore pass startup profiling, serve ordinary traffic for hours, and fail the first time the arriving prompts are long enough to build a full-sized batch.
- The reason the served context length is not the lever is that the prompts were always inside it. What changed is how many prefill tokens the scheduler is willing to put into one step, which is a throughput setting rather than a capacity one.
The fix and how to prevent it
Searchable error signature
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 3.83 GiB. GPU 0 has a total capacity of 79.15 GiB of which 2.11 GiB is free.
INFO: Chunked prefill is enabled with max_num_batched_tokens=16384.
WARNING: If out-of-memory occurs during cudagraph capture, consider decreasing gpu_memory_utilization or switching to eager mode.Use this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Why the recommended fix works
Lowering the token budget shrinks the largest step the scheduler can build, which lowers the activation peak that collides with the reserved cache. It costs time to first token rather than correctness, and it targets the quantity that actually changed, unlike lowering the served context length, which does nothing when the prompts were already inside it.
Best practices by model family
| Model / Stack | Recommendation | Notes |
|---|---|---|
| Runtime out-of-memory after a throughput change | Lower --max-num-batched-tokens to 2048 to 4096 | This is the setting that enlarged the step and its activation peak. |
| Large model or long-context serving | --max-num-seqs 128 to 256 | Bounds how many sequences contribute to one batch. |
| Throughput-oriented small model on a large GPU | A budget above 8192, validated at long-prompt load | Better time to first token, but only safe if tested at the prompt lengths that fill it. |
| Failure during graph capture | --enforce-eager | Removes the memory reserved for CUDA graphs from the peak. |
With the fix vs without the fix
| Dimension | With the fix | Without the fix |
|---|---|---|
| When the failure appears | During serving, on long prompts | Expected at startup, like a cache sizing error |
| Which budget is responsible | The per-step batched token limit | Assumed to be gpu_memory_utilization |
| Effect of lowering max_model_len | No help when prompts were already inside it | Assumed to reduce memory in every case |
Diagnostic note
“There is no closed-form value for this budget, and looking for one wastes time. The working method upstream and in practice is to raise it until the deployment fails at representative load and then step back one notch, because the limit depends on the model, the sequence mix and the card. What matters is that the load used for that experiment contains the longest prompts the service will really see; tuning against short prompts produces a value that fails the first time a long one arrives.”
Visual fingerprint
reserved by gpu_memory_utilization [ weights ][ KV cache pool ]
needed per step by the token budget [ activations for max_num_batched_tokens ]
^
grows when the budget is raised, until it
collides with the reserved pool at runtimeDiagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionvLLM errors in context
vLLM failures cross request, model, KV cache, worker, kernel and host boundaries. The hub compares common literal signatures and the first control that separates their causal owners.
Compare every vllm error side by sideRelated failures to investigate next
Frequently asked questions
Questions engineers and on-call staff commonly ask about this failure.
Why did it start fine and fail later?
Should I lower gpu_memory_utilization?
Is there a formula for max_num_batched_tokens?
Does lowering max_model_len help?
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.