TensorRT-LLM refuses a quantized checkpoint because its weight mapper does not recognise the export layout
TensorRT-LLM loads a checkpoint through a per-model weight mapper that expects tensor names and groupings in a particular layout. A quantized export written by a different tool, or by a different version of the same tool, can be entirely valid and still be rejected, because the mapper has no rule for the names it contains.
The file is fine and the reader has no rule for it. Load it with the tool that produced it to prove that, then align the quantizer version with the TensorRT-LLM version rather than editing the checkpoint.
- Symptom
[TRT-LLM] [E] qwen3_5_weight_mapper rejects ModelOpt W4A16_AWQ checkpoints- Root cause
- Loading is a translation step, not a copy. Each supported architecture has a mapper that knows which exported tensor becomes which internal parameter, including how quantized weights, scales and zero points are grouped. That knowledge is written per scheme, so a scheme it was not written for has no route through.
- Recommended fix
- Establish first that the file is intact by loading it with the tool that produced it. A successful load there rules out corruption and localises the problem to the mapper.
- How Denpex helps
- Denpex matches TensorRT-LLM refuses a quantized checkpoint because its weight mapper does not recognise the export layout across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
What this failure is
A checkpoint load failure caused by a mismatch between the tensor layout a quantized export was written in and the layout the target architecture's weight mapper was written to read, with no corruption of the file involved.
Is this what broke your run? Paste your log.
You're reading about TensorRT-LLM refuses a quantized checkpoint because its weight mapper does not recognise the export layout. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.
Why it happens (the mechanism)
A quantized checkpoint is not one tensor per parameter. It carries packed weights alongside scales and zero points whose names and grouping are chosen by the exporting tool, and the mapper on the reading side hard-codes the arrangement it was written against. Two correct programs can therefore disagree about a correct file, and they do so most often when their versions have drifted apart.
What you'll observe
- A checkpoint that other tools load without complaint is refused here
- The rejection names a mapper or a key rather than saying what layout was expected
- Re-exporting with different settings sometimes works, without making clear which setting mattered
- The same model in an unquantized form loads normally
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| AssertionError: Expected packed qkv scale tensor with leading dim 64, got torch.Size([8192, 20]) | Loading is a translation step, not a copy. Each supported architecture has a mapper that knows which exported tensor becomes which internal parameter, including how quantized weights, scales and zero points are grouped. That knowledge is written per scheme, so a scheme it was not written for has no route through. |
| assert tensor.shape[0] == expected_total_blocks raised in _split_qkv_scale_tensor | Quantized exports carry additional tensors that unquantized ones do not, and their names differ between producers and between versions of one producer. The mapper is therefore coupled to the exporting tool in a way the checkpoint format alone does not express, which is why a file can be internally consistent and still unloadable. |
| Traceback ending in tensorrt_llm/_torch/models/checkpoints/hf/qwen3_5_weight_mapper.py | The rejection is not evidence of corruption. Nothing about the bytes is wrong; the reader simply has no rule for the layout it found, which is why other tools accept the identical file. |
| Checkpoint load aborts inside the architecture weight mapper before any engine build begins | Loading is a translation step, not a copy. Each supported architecture has a mapper that knows which exported tensor becomes which internal parameter, including how quantized weights, scales and zero points are grouped. That knowledge is written per scheme, so a scheme it was not written for has no route through. |
| A weight mapper rejecting ModelOpt W4A16_AWQ checkpoints during load | Quantized exports carry additional tensors that unquantized ones do not, and their names differ between producers and between versions of one producer. The mapper is therefore coupled to the exporting tool in a way the checkpoint format alone does not express, which is why a file can be internally consistent and still unloadable. |
| Errors naming a model-specific weight mapper rather than a file or a checksum | The rejection is not evidence of corruption. Nothing about the bytes is wrong; the reader simply has no rule for the layout it found, which is why other tools accept the identical file. |
Which systems are affected
- ModelOpt exports in W4A16 AWQ and neighbouring quantization schemes
- Model families whose mapper was written against one export tool's naming
- Version skew between the quantizer that wrote the checkpoint and the TensorRT-LLM that reads it
- Custom or community quantizations reusing an architecture name without matching its expected layout
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Load the checkpoint with the tool that wrote it. Success confirms the bytes are sound and the problem is the reader's expectations.
- ✓List the tensor names in the export and compare a sample against those the mapper looks for. A consistent difference in prefix or grouping identifies the mismatch.
- ✓Load the unquantized variant of the same model. Success there shows the architecture is supported and narrows the fault to the quantized layout.
Root cause
- Loading is a translation step, not a copy. Each supported architecture has a mapper that knows which exported tensor becomes which internal parameter, including how quantized weights, scales and zero points are grouped. That knowledge is written per scheme, so a scheme it was not written for has no route through.
- Quantized exports carry additional tensors that unquantized ones do not, and their names differ between producers and between versions of one producer. The mapper is therefore coupled to the exporting tool in a way the checkpoint format alone does not express, which is why a file can be internally consistent and still unloadable.
- The rejection is not evidence of corruption. Nothing about the bytes is wrong; the reader simply has no rule for the layout it found, which is why other tools accept the identical file.
The fix and how to prevent it
Searchable error signature
[TRT-LLM] [E] qwen3_5_weight_mapper rejects ModelOpt W4A16_AWQ checkpoints
[TRT-LLM] [I] Loading weights via the architecture weight mapperUse this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Why the recommended fix works
Aligning the producer and the runtime removes the drift that the mapper is sensitive to, so the names it looks for are the names present. Re-exporting to a documented scheme works for the same reason from the other side. Comparing tensor names turns the disagreement into a specific, visible difference rather than a trial of settings.
Best practices by model family
| Model / Stack | Recommendation | Notes |
|---|---|---|
| ModelOpt export refused | Match the ModelOpt version to the TensorRT-LLM release | Naming a mapper expects is set by the producer it was written against. |
| Scheme close to a supported one | Re-export to the documented scheme for the architecture | Neighbouring schemes differ in layout while sounding interchangeable. |
| Suspected file damage | Load with the producing tool first | A clean load rules out corruption in one step. |
| Custom or community quantization | Compare tensor names against the mapper's expectations | A shared architecture name does not imply a shared layout. |
With the fix vs without the fix
| Dimension | With the fix | Without the fix |
|---|---|---|
| What the rejection means | The reader has no rule for this layout | Read as the checkpoint being corrupt |
| Where compatibility lives | Between the quantizer version and the runtime version | Assumed to be a property of the file alone |
| Why other tools load it | They have their own mappers | Taken as proof the runtime is at fault |
Diagnostic note
“The word rejects invites a hunt for a damaged file, and that hunt is always fruitless here. The strongest early move is to load the checkpoint with its own producer: it takes a minute, and a clean load converts the question from is this file broken into which of these two tools is out of date. Skipping that step is how teams end up re-downloading weights repeatedly and re-quantizing at random.”
Visual fingerprint
quantizer export architecture weight mapper
layer.qweight expects layer.weight_packed
layer.scales --> expects layer.weight_scale
layer.zeros expects layer.weight_zero_point
|
no rule for these names -> rejectedDiagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionRelated failures to investigate next
Frequently asked questions
Questions engineers and on-call staff commonly ask about this failure.
Is my checkpoint corrupt?
Why does another framework load it fine?
Which version should I change?
Can I rename the tensors by hand?
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.