DeepGEMM rejects FP8 block-scaled weights with Unknown SF transformation on an unsupported architecture
A block-scaled FP8 checkpoint fails while its weights are being prepared, before the server ever accepts a request. The runtime selected a specialised matrix-multiply library for the detected architecture, and that library does not implement the scale-factor layout this class of device requires, so it refuses the transformation outright.
The runtime picked a kernel library your GPU is not actually supported by, because the support check accepts a whole capability family. Use a general-purpose quantised path, a per-tensor checkpoint, or a data-centre part of the same generation.
- Symptom
RuntimeError: Assertion error (/workspace/.deps/deepgemm-src/csrc/apis/layout.hpp:60): Unknown SF transformation- Root cause
- Block scaling stores a grid of scale factors alongside the weights, and a kernel consumes them in a device-specific arrangement. Transforming the stored grid into that arrangement is a per-architecture routine, and a device family the library has not implemented has no routine to call, hence a refusal naming the transformation rather than the device. The runtime reached that library because its support check groups compute capability into families and accepts the whole family.
- Recommended fix
- Select a general-purpose quantised matrix-multiply path instead of the specialised library, which trades throughput for a kernel that covers the device.
- How Denpex helps
- Denpex matches DeepGEMM rejects FP8 block-scaled weights with Unknown SF transformation on an unsupported architecture across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
What this failure is
A load-time failure in which a serving runtime selects a specialised FP8 matrix-multiply library for a device inside an accepted compute-capability family, and the library has no scale-factor layout transformation implemented for that architecture.
Is this what broke your run? Paste your log.
You're reading about DeepGEMM rejects FP8 block-scaled weights with Unknown SF transformation on an unsupported architecture. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.
Why it happens (the mechanism)
Kernel selection has to decide from something, and compute capability is the number available at runtime. Grouping it into families is a reasonable approximation right up until a family contains parts with genuinely different capabilities, which is what happens when a generation spans consumer and data-centre silicon. The runtime then admits a device to a fast path whose library never claimed it, and the refusal comes from the library rather than from the check that should have made it.
What you'll observe
- The failure is at weight preparation, so nothing serves and there is no partial capability to fall back to
- The message is from a kernel library's internals and names a layout concept, not a device or a model
- The same checkpoint loads correctly on data-centre parts of the same generation
- Both workers fail identically, which makes it look like a model problem rather than a hardware-support one
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| RuntimeError: Assertion error naming a layout header and the text Unknown SF transformation | Block scaling stores a grid of scale factors alongside the weights, and a kernel consumes them in a device-specific arrangement. Transforming the stored grid into that arrangement is a per-architecture routine, and a device family the library has not implemented has no routine to call, hence a refusal naming the transformation rather than the device. |
| The failure raised from process_weights_after_loading, during load and before any inference request | The runtime reached that library because its support check groups compute capability into families and accepts the whole family. Consumer and data-centre parts of a generation share a capability family while differing in the tensor-core and memory features these kernels are written against, so the check admits a device the kernel library never claimed. |
| A call chain descending from the quantisation scheme into a scaled matrix-multiply kernel and then into scale-factor layout transformation | It fails during preparation because that is when the scales are rearranged into the layout the kernel expects. This is fortunate: the alternative would be selecting an unsupported path at request time, after the service had reported itself healthy. |
| Every tensor-parallel worker failing at the same point, since they all prepare the same weights | Block scaling stores a grid of scale factors alongside the weights, and a kernel consumes them in a device-specific arrangement. Transforming the stored grid into that arrangement is a per-architecture routine, and a device family the library has not implemented has no routine to call, hence a refusal naming the transformation rather than the device. |
Which systems are affected
- Block-scaled FP8 checkpoints, where each block of weights carries its own scale rather than one scale per tensor
- Consumer Blackwell parts reporting compute capability 12.0, which the runtime groups with data-centre Blackwell for kernel selection
- Serving runtimes that choose a specialised GEMM library from a compute-capability family rather than from a tested device list
- Tensor-parallel deployments, where the identical preparation runs on every worker and fails on all of them
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Check the reported compute capability of the device against the parts the kernel library documents. A capability inside an accepted family but outside the implemented list is the whole condition.
- ✓Load a per-tensor quantised version of the same model. Succeeding where the block-scaled one failed confirms the scale layout rather than the checkpoint.
- ✓Force a general-purpose quantised path and reload. Starting successfully identifies kernel selection as the decision that failed.
Root cause
- Block scaling stores a grid of scale factors alongside the weights, and a kernel consumes them in a device-specific arrangement. Transforming the stored grid into that arrangement is a per-architecture routine, and a device family the library has not implemented has no routine to call, hence a refusal naming the transformation rather than the device.
- The runtime reached that library because its support check groups compute capability into families and accepts the whole family. Consumer and data-centre parts of a generation share a capability family while differing in the tensor-core and memory features these kernels are written against, so the check admits a device the kernel library never claimed.
- It fails during preparation because that is when the scales are rearranged into the layout the kernel expects. This is fortunate: the alternative would be selecting an unsupported path at request time, after the service had reported itself healthy.
The fix and how to prevent it
Searchable error signature
RuntimeError: Assertion error (/workspace/.deps/deepgemm-src/csrc/apis/layout.hpp:60): Unknown SF transformation
compressed_tensors_w8a8_fp8.py:169 process_weights_after_loading
-> kernels/linear/scaled_mm/deep_gemm.py:96 process_weights_after_loading
-> quantization/utils/fp8_utils.py:1077 deepgemm_post_process_weight_scale_block
-> utils/deep_gemm.py:494 transform_sf_into_required_layoutUse this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Why the recommended fix works
A general-purpose quantised kernel is written against the baseline features of the architecture rather than against a specific product's tensor cores, so it covers the device at a lower throughput. Moving to per-tensor scaling removes the requirement entirely, because a single scale per tensor needs no layout transformation. Both replace an unimplemented routine with one that exists, which is the only thing standing between the checkpoint and a running server.
Best practices by model family
| Model / Stack | Recommendation | Notes |
|---|---|---|
| Consumer Blackwell, block-scaled FP8 | General-purpose quantised path | Lower throughput, but the kernel exists. |
| Consumer part, any FP8 | Per-tensor scaled checkpoint | No scale-factor layout transformation required. |
| Data-centre part | Keep the specialised library | This is the hardware the fast path targets. |
| Memory permits | Serve in half precision | Avoids quantised kernel selection altogether. |
With the fix vs without the fix
| Dimension | With the fix | Without the fix |
|---|---|---|
| What decided the kernel | A compute-capability family | Assumed to be a tested device list |
| Where it fails | Weight preparation, before serving | Expected at request time |
| Is the checkpoint at fault | No, the kernel path is | Investigated as a bad download |
Diagnostic note
“Worth internalising: a support check that reasons about a capability family is a guess about hardware, and the guess is wrong exactly where a generation spans market segments. When a quantised model refuses to load on a consumer card of an otherwise supported generation, look at the kernel selection logic before you look at the checkpoint. The checkpoint is usually fine and will load the moment a different path is chosen.”
Visual fingerprint
device reports capability 12.0
|
v
runtime: family 100 or 120 -> select specialised FP8 library
|
v
library: transform block scales into required layout
|
v
no routine for this architecture -> Unknown SF transformationDiagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionFrequently asked questions
Questions engineers and on-call staff commonly ask about this failure.
Is my checkpoint corrupt?
My GPU is the right generation. Why is it unsupported?
What does the fallback cost?
Would a newer runtime fix it?
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.