Skip to content

vLLM rejects a request because the prompt plus completion exceeds the served context length

vLLM returns an error for a request whose prompt tokens plus requested completion tokens exceed the context window the server was started with. The served window is not always the one advertised on the model card: it is taken from the model config and can be capped further when KV cache memory is short.

Quick answer

Add the prompt length to max_tokens and compare that sum, not the prompt alone, against the served window. Then set --max-model-len explicitly rather than trusting the value vLLM infers from the model config.

Symptom
ValueError: This model's maximum context length is 4096 tokens, however you requested 5120 tokens (4608 in your prompt; 512 for the completion). Reduce max_model_len or the prompt length.
Root cause
The limit is applied to prompt tokens plus requested completion tokens, not to the prompt alone. A prompt that fits on its own still fails when max_tokens is added to it, which is why reducing the requested completion sometimes clears the error without touching the input. vLLM derives the served window from the model configuration, reading max_position_embeddings.
Recommended fix
Set --max-model-len explicitly at launch to the length you intend to serve, rather than relying on the value inferred from the model config.
How Denpex helps
Denpex matches vLLM rejects a request because the prompt plus completion exceeds the served context length across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
Memory#vllm#max-model-len#context-length#inference-serving#max-position-embeddings#request-rejection

What this failure is

A request-time rejection in which vLLM refuses a completion because the prompt tokens plus the requested completion tokens exceed the context window the server is actually serving, which is often shorter than the window the model card advertises.

Live diagnosis, no signup

Is this what broke your run? Paste your log.

You're reading about vLLM rejects a request because the prompt plus completion exceeds the served context length. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.

training_logs.txt
No log to hand? Try one:

3 free diagnoses/day

Why it happens (the mechanism)

Two independent things shorten the served window. The model config's max_position_embeddings, not the model card, is what vLLM reads, and repositories exist where the card advertises a long context while the config declares a short one. Separately, the engine caps the window at startup when the KV cache cannot hold the full length within the memory budget. Either produces a served window smaller than expected, and every request is then judged against that smaller number, with the completion counted alongside the prompt.

What you'll observe

  • A request is rejected although the model card advertises a much longer context
  • The same prompt succeeds against a different deployment of the same model
  • Shortening the prompt clears the error even though the completion length was never changed
  • Raising the served context length triggers a different startup error about KV cache capacity

Common symptoms and what they mean

SymptomWhy it happens
ValueError: This model's maximum context length is 4096 tokens, however you requested 5120 tokens. Reduce max_model_len or the prompt length.The limit is applied to prompt tokens plus requested completion tokens, not to the prompt alone. A prompt that fits on its own still fails when max_tokens is added to it, which is why reducing the requested completion sometimes clears the error without touching the input.
The model's max seq len is larger than the maximum number of tokens that can be stored in KV cache. Try increasing gpu_memory_utilization or decreasing max_model_len when initializing the engine.vLLM derives the served window from the model configuration, reading max_position_embeddings. A repository whose card advertises a long context while its config.json declares a shorter max_position_embeddings will serve the shorter one, and the error names that shorter number.
A served context length that matches max_position_embeddings in the model config rather than the length the model card advertisesThe served window can also be capped below the configured one at startup when the KV cache cannot hold the full length within the memory budget. In that case the engine logs the cap and every later request is judged against the reduced number.

Which systems are affected

  • vLLM OpenAI-compatible server endpoints, which reproduce the OpenAI error wording
  • Models whose config.json carries a max_position_embeddings smaller than the advertised context
  • Long-context deployments where KV cache memory, not the model, is the binding limit
  • Clients that hardcode max_tokens instead of deriving it from the remaining budget

How to confirm this is the problem

Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.

  • ✓Add the prompt token count to the requested max_tokens and compare the sum against the number in the error; the error reports both halves separately for this purpose.
  • ✓Read max_position_embeddings from the model's config.json and compare it against the length the engine logged at startup.
  • ✓Check the startup log for a warning that the sequence length was capped by KV cache capacity, which explains a served window shorter than the one requested.

Root cause

  • The limit is applied to prompt tokens plus requested completion tokens, not to the prompt alone. A prompt that fits on its own still fails when max_tokens is added to it, which is why reducing the requested completion sometimes clears the error without touching the input.
  • vLLM derives the served window from the model configuration, reading max_position_embeddings. A repository whose card advertises a long context while its config.json declares a shorter max_position_embeddings will serve the shorter one, and the error names that shorter number.
  • The served window can also be capped below the configured one at startup when the KV cache cannot hold the full length within the memory budget. In that case the engine logs the cap and every later request is judged against the reduced number.

The fix and how to prevent it

Searchable error signature

search key
ValueError: This model's maximum context length is 4096 tokens, however you requested 5120 tokens (4608 in your prompt; 512 for the completion). Reduce max_model_len or the prompt length.
WARNING: The model's max seq len (32768) is larger than the maximum number of tokens that can be stored in KV cache (15424).

Use this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.

The fix and the prevention pattern

The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.

Sign up free. Unlock the full analysis

No credit card. Daily allowance follows verified trust tier. Instant access.

Why the recommended fix works

Setting --max-model-len explicitly removes the inference step entirely, so the served window becomes a value you chose rather than one derived from a config file you may not have read. Deriving max_tokens on the client from the remaining budget fixes the other half, because the limit applies to the sum and a hardcoded completion length can exceed it no matter how short the prompt is.

Best practices by model family

Model / StackRecommendationNotes
Any production deploymentPass --max-model-len explicitlyRemoves dependence on max_position_embeddings and on release-to-release inference changes.
Model card and config disagreeTrust config.json, then override deliberatelyvLLM reads max_position_embeddings; the card is a claim, not a setting.
Long context on limited VRAMRaise --gpu-memory-utilization or quantize the KV cacheWhen the cache is the binding limit, memory rather than the model decides the window.
Client integrationsCompute max_tokens from the remaining budgetThe limit is on prompt plus completion, so a fixed max_tokens will eventually exceed it.

With the fix vs without the fix

DimensionWith the fixWithout the fix
What the limit applies toPrompt tokens plus requested completion tokensAssumed to apply to the prompt alone
Source of the served windowmax_position_embeddings, possibly capped by KV cache capacityAssumed to be the context on the model card
Effect of raising the windowMay trigger a KV cache capacity error at startupAssumed to be free

Diagnostic note

“Overriding the length does not grant long-context quality. Deployments have been reported emitting degraded or random output well before the overridden limit, so a model forced to 128k because the card claims 128k can produce confident nonsense at 32k. Treat an override as a serving-capacity decision that still needs evaluation at the lengths you intend to use.”

Visual fingerprint

The limit is on the sum, not the prompt
served window (max_model_len)  |==============================| 4096
prompt                         |==========================|     3600
requested completion           |                          |==|  1024
total requested                |=============================>| 4624  <-- rejected
A 3600-token prompt fits inside a 4096-token window on its own, but adding a 1024-token completion request brings the total to 4624 and the request is rejected. Reducing max_tokens alone clears it without shortening the prompt.

Diagnose this failure in VS Code

Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.

Install the free VS Code extension

vLLM errors in context

vLLM failures cross request, model, KV cache, worker, kernel and host boundaries. The hub compares common literal signatures and the first control that separates their causal owners.

Compare every vllm error side by side

Frequently asked questions

Questions engineers and on-call staff commonly ask about this failure.

The model supports 128k context. Why does the error say 8192?
vLLM reads max_position_embeddings from config.json. Some repositories advertise a long context on the card while declaring a shorter max_position_embeddings, and the shorter value is what gets served.
My prompt is shorter than the limit, so why was it rejected?
The limit applies to prompt tokens plus requested completion tokens. A prompt that fits alone can still exceed the window once max_tokens is added.
I raised --max-model-len and now the server will not start.
The KV cache cannot hold the longer window within the memory budget. Raise --gpu-memory-utilization, quantize the KV cache, or add tensor parallelism.
Is overriding the context length safe?
It changes what the server accepts, not what the model handles well. Evaluate output quality at the lengths you intend to serve, because degradation well before the limit has been reported.

Don't just read the fix, diagnose your run

The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.