TensorRT-LLM FP8 support depends on GPU architecture, backend and model
An FP8 architecture rejection requires checking the selected TensorRT-LLM release, backend, quantization recipe and GPU. NVIDIA's reviewed support matrix includes Ada SM89 and Hopper SM90 FP8, but does not list native FP8 for Ampere SM80 or SM86.
Check the exact GPU and TensorRT-LLM support matrix before changing flags. Ada SM89 (compute capability 8.9) is included in NVIDIA's reviewed FP8 precision table; a strict greater-than-8.9 rule would incorrectly exclude it. Ampere is not listed for native FP8 in that table. Hardware support alone does not guarantee that a particular model, backend, kernel or quantization recipe supports FP8. Use a documented supported format for that combination and validate quality and memory use.
- Symptom
ERROR: FP8 quantization is not supported on this GPU architecture- Root cause
- The requested precision path is not supported by the detected architecture and software combination. Model or backend support can be narrower than hardware capability. A mismatched or reused engine artifact can require a separate compatibility investigation.
- Recommended fix
- Record the GPU model, compute capability, TensorRT-LLM version, backend, model and complete quantization configuration.
- How Denpex helps
- Denpex matches TensorRT-LLM FP8 support depends on GPU architecture, backend and model across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
What this failure is
An FP8 architecture rejection means the selected build or runtime path cannot use the requested FP8 configuration on the detected target. It is not evidence that every FP8-related representation or every backend has the same support boundary.
Is this what broke your run? Paste your log.
You're reading about TensorRT-LLM FP8 support depends on GPU architecture, backend and model. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.
Why it happens (the mechanism)
TensorRT-LLM support is a combination of GPU architecture, software release, backend, model and precision path. NVIDIA's reviewed support matrix lists FP8 for Ada SM89 and Hopper SM90, and lists other precisions for Ampere SM80 and SM86. The same documentation warns that quantized formats are not implemented for every model. A performance-guide sentence uses a strict greater-than-8.9 threshold while naming Ada; use the explicit architecture table rather than repeating that inconsistent inequality.
What you'll observe
- A requested FP8 path is rejected on the detected GPU
- A general hardware capability claim does not match the selected model or backend
- The build environment and deployment environment use different GPU architectures or software versions
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| An error mentions FP8 and an unsupported GPU architecture or compute capability | The requested precision path is not supported by the detected architecture and software combination |
| The selected precision does not appear in the applicable architecture and model support table | Model or backend support can be narrower than hardware capability |
| A configuration works in one environment but fails in another | A mismatched or reused engine artifact can require a separate compatibility investigation |
Which systems are affected
- TensorRT-LLM builds or inference paths requesting FP8
- Ampere deployments requesting a native FP8 path not listed for SM80 or SM86
- Ada, Hopper and newer deployments whose selected model, backend or recipe has additional restrictions
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Read the visible GPU compute capability and installed package versions
- ✓Match the selected model and precision recipe to the installed release's support table
- ✓Capture the complete error and identify whether rejection occurs in conversion, engine build or inference
- ✓Compare build and deployment environments without assuming a cross-build workflow is supported
Root cause
- The requested precision path is not supported by the detected architecture and software combination
- Model or backend support can be narrower than hardware capability
- A mismatched or reused engine artifact can require a separate compatibility investigation
The fix and how to prevent it
Searchable error signature
ERROR: FP8 quantization is not supported on this GPU architectureUse this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Why the recommended fix works
Selecting a supported architecture, model and precision path removes a compatibility mismatch. Measuring the resulting model separately avoids replacing an unsupported format with an unverified accuracy or performance claim.
Code examples
# Read-only environment inventory. Run in the environment that reports the error.
python - <<'PY'
import importlib.metadata as metadata
import torch
print('PyTorch:', torch.__version__, 'CUDA build:', torch.version.cuda)
try:
print('TensorRT-LLM:', metadata.version('tensorrt-llm'))
except metadata.PackageNotFoundError:
print('TensorRT-LLM package metadata not found in this environment')
for index in range(torch.cuda.device_count()):
print(index, torch.cuda.get_device_name(index), torch.cuda.get_device_capability(index))
PYAdapt the snippet to your framework. The same pattern holds for PyTorch Lightning, Hugging Face Trainer, DeepSpeed, Megatron-LM, and vLLM training wrappers. Where the wrapper exposes a config flag (for examplelr_scheduler_type in Trainer), prefer the flag over the imperative API to keep the schedule declarative and reproducible.
Best practices by model family
| Model / Stack | Recommendation | Notes |
|---|---|---|
| Ada SM89 | Check the exact model and backend FP8 path | Compute capability 8.9 is included in the reviewed architecture precision table. |
| Hopper SM90 | Validate the selected FP8 recipe | Hardware capability does not certify every model or kernel. |
| Ampere SM80 or SM86 | Use a precision path explicitly supported for the deployment | The reviewed matrix does not list native FP8 for these architectures. |
With the fix vs without the fix
| Dimension | With the fix | Without the fix |
|---|---|---|
| Compatibility decision | Pinned release, backend, model and architecture evidence | One universal compute-capability inequality |
| Alternative precision | Supported recipe plus quality and resource measurements | An assumption of equivalent throughput or memory savings |
Diagnostic note
“Documentation review, September 27, 2026. The reviewed unversioned TensorRT-LLM documentation identifies itself as 1.1.0rc5. This is not a claim of customer incident experience or validation of all later releases. The example is a representative error fragment, not a captured incident. The architecture table is explicit about Ada SM89; an inconsistent greater-than-8.9 sentence in the performance guide is not repeated as a universal requirement.”
Visual fingerprint
GPU + release + backend + model + recipe -> supported path -> measured validation
Diagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionFrequently asked questions
Questions engineers and on-call staff commonly ask about this failure.
Does compute capability 8.9 exclude Ada from FP8 support?
Can a flag add native FP8 Tensor Core support to an A100?
Will switching to INT4 or INT8 preserve the same accuracy and speed?
Does successful execution on H100 prove the same configuration works on Ada?
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.