Skip to content

TensorRT-LLM fails at model init because the FlashInfer FMHA kernel has no build for the GPU architecture

TensorRT-LLM delegates fused attention to FlashInfer, whose FMHA runner is compiled for a specific set of architectures. On a device outside that set the runner raises an unsupported-architecture error from its own header, so the model finishes loading and then fails to serve.

Quick answer

The bundled FlashInfer build has no kernel compiled for this compute capability. Move to an image whose kernels cover the device, or serve on a generation it does cover; no TensorRT-LLM flag adds the missing object code.

Symptom
RuntimeError: Error in function 'TllmGenFmhaRunner' at /usr/local/lib/python3.12/dist-packages/flashinfer/data/include/flashinfer/trtllm/fmha/fmhaRunner.cuh:37: Unsupported architecture
Root cause
Fused attention is not written per model; it is a prebuilt kernel selected at runtime. The build carries object code for an enumerated list of architectures, and a device outside that list has nothing to dispatch to, so the runner refuses rather than falling back. The list is a property of the bundled FlashInfer build rather than of TensorRT-LLM or of the checkpoint, which is why the error names a header inside FlashInfer and why nothing in the invocation hints at it.
Recommended fix
Read the compute capability of the target device and compare it against the architectures the bundled kernel was built for, rather than assuming a current release covers current hardware.
How Denpex helps
Denpex matches TensorRT-LLM fails at model init because the FlashInfer FMHA kernel has no build for the GPU architecture across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
Environment#tensorrt-llm#flashinfer#fmha#unsupported-architecture#sm120#blackwell

What this failure is

A runtime rejection in which the FlashInfer fused multi-head attention runner bundled with TensorRT-LLM finds no compiled kernel for the target GPU architecture and raises rather than falling back, after model initialisation has already succeeded.

Live diagnosis, no signup

Is this what broke your run? Paste your log.

You're reading about TensorRT-LLM fails at model init because the FlashInfer FMHA kernel has no build for the GPU architecture. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.

training_logs.txt
No log to hand? Try one:

3 free diagnoses/day

Why it happens (the mechanism)

Attention kernels are compiled ahead of time for an enumerated list of architectures and dispatched at runtime. Dispatch either finds an entry for the device or it does not, and there is no generic path to degrade to, so the runner raises from inside its own header. Because the enumeration belongs to the bundled kernel library rather than to the framework, nothing in the TensorRT-LLM invocation reflects it.

What you'll observe

  • The container release is current and the model is supported, yet initialisation fails on this card
  • The error is raised from a FlashInfer header path, which does not appear in the TensorRT-LLM command that was run
  • Both a small inference card and a new flagship can hit it, which makes the pattern look inconsistent
  • Model loading reports success first, so the failure looks like a serving problem rather than a support one

Common symptoms and what they mean

SymptomWhy it happens
RuntimeError: Error in function 'TllmGenFmhaRunner' at flashinfer/trtllm/fmha/fmhaRunner.cuh: Unsupported architectureFused attention is not written per model; it is a prebuilt kernel selected at runtime. The build carries object code for an enumerated list of architectures, and a device outside that list has nothing to dispatch to, so the runner refuses rather than falling back.
Failure occurring after a model init total timing line, not during weight downloadThe list is a property of the bundled FlashInfer build rather than of TensorRT-LLM or of the checkpoint, which is why the error names a header inside FlashInfer and why nothing in the invocation hints at it.
The same checkpoint serving normally on a different GPU generation with an unchanged commandCoverage is not monotonic with newness. Hardware released after the image predates nothing it can use, and modest inference parts are sometimes left out of a build aimed at data-centre silicon, so both the newest and the smallest devices fail while the middle of the range works.
Quantization configuration naming NVFP4 with modelopt as the producerFused attention is not written per model; it is a prebuilt kernel selected at runtime. The build carries object code for an enumerated list of architectures, and a device outside that list has nothing to dispatch to, so the runner refuses rather than falling back.

Which systems are affected

  • GPUs whose compute capability is outside the set the bundled FlashInfer FMHA was compiled for, in both directions
  • Newer silicon released after the container image, such as SM120 Blackwell workstation parts
  • Smaller inference cards such as the L4, which are sometimes omitted from a kernel's build list
  • Release-candidate TensorRT-LLM containers, whose bundled kernel coverage moves between builds

How to confirm this is the problem

Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.

  • ✓Query the device compute capability and check it against the architecture list the installed FlashInfer build advertises. A device absent from that list is the whole explanation.
  • ✓Run the identical image and command on a different GPU generation. Success there separates a kernel coverage gap from a problem with the model or the configuration.
  • ✓Confirm the failure occurs after the model init timing line. A rejection at that point is kernel dispatch rather than weight loading.

Root cause

  • Fused attention is not written per model; it is a prebuilt kernel selected at runtime. The build carries object code for an enumerated list of architectures, and a device outside that list has nothing to dispatch to, so the runner refuses rather than falling back.
  • The list is a property of the bundled FlashInfer build rather than of TensorRT-LLM or of the checkpoint, which is why the error names a header inside FlashInfer and why nothing in the invocation hints at it.
  • Coverage is not monotonic with newness. Hardware released after the image predates nothing it can use, and modest inference parts are sometimes left out of a build aimed at data-centre silicon, so both the newest and the smallest devices fail while the middle of the range works.

The fix and how to prevent it

Searchable error signature

search key
RuntimeError: Error in function 'TllmGenFmhaRunner' at /usr/local/lib/python3.12/dist-packages/flashinfer/data/include/flashinfer/trtllm/fmha/fmhaRunner.cuh:37: Unsupported architecture
Model init total -- 21.36s
[TRT-LLM] [W] [quantize] Inline 'config.json.quantization_config' diverges from 'hf_quant_config.json'

Use this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.

The fix and the prevention pattern

The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.

Sign up free. Unlock the full analysis

No credit card. Daily allowance follows verified trust tier. Instant access.

Why the recommended fix works

Changing to a build that enumerates the architecture supplies the object code dispatch was looking for, which is the only thing missing. Selecting a different attention provider works for the same reason: it moves the request to a kernel set that does cover the device. Neither is a tuning change, because nothing about the model or its configuration was wrong.

Best practices by model family

Model / StackRecommendationNotes
Newly released GPU generationMove to a later container imageThe kernel needs the architecture compiled in; no flag substitutes.
Smaller inference card such as the L4Check the bundled kernel's architecture list explicitlyModest parts are sometimes omitted from data-centre-oriented builds.
Framework allows a backend choiceSelect an attention provider that covers the deviceThe rejection comes from one kernel provider, not from attention itself.
Qualifying new hardwareRun a one-request serve check, not just a load checkThe failure appears after model init reports success.

With the fix vs without the fix

DimensionWith the fixWithout the fix
What decides supportThe architecture list compiled into the bundled kernelAssumed to be the TensorRT-LLM version
Direction of the gapBoth newer and smaller devices can be missingRead as the device being too old
Where it failsAfter model init, at kernel dispatchExpected during weight loading

Diagnostic note

“It is tempting to read unsupported architecture as too old. Both reported cases contradict that: one is a Blackwell workstation card newer than the image, the other a small inference part quietly absent from a build aimed at larger silicon. The useful question is not whether the device is modern but whether this particular kernel build enumerates it, and the answer changes between release candidates of the same version.”

Visual fingerprint

Dispatch either finds the architecture or refuses
  bundled FlashInfer FMHA build
    compiled for:  SM90  SM100        <- enumerated at build time
                      |
  target device  SM120  --> not in the list --> Unsupported architecture
  target device  SM89   --> not in the list --> Unsupported architecture
  target device  SM90   --> dispatch succeeds
The kernel carries object code only for the architectures enumerated when it was built. A device outside that list has nothing to dispatch to and the runner raises, which is why both a newer part and a smaller part can fail while the middle of the range serves normally.

Diagnose this failure in VS Code

Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.

Install the free VS Code extension

Frequently asked questions

Questions engineers and on-call staff commonly ask about this failure.

My GPU is brand new. How can it be unsupported?
The kernel was compiled before the device existed, so its architecture list does not include it. Newness is the cause rather than a defence.
Is there a flag to force it?
No. The object code for the architecture is either present in the build or it is not, and no runtime option compiles it.
Why did the model load successfully first?
Weight loading does not dispatch attention kernels. The rejection happens later, when the fused attention runner is first asked for an implementation.
Does this mean TensorRT-LLM does not support my card?
Not necessarily. It means this bundled kernel build does not. A later image, or a different attention backend, may cover it.

Don't just read the fix, diagnose your run

The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.