Update · 4 Oct 2026 — Two more corrections, both in the corrections section below. NVIDIA's NVFP4 checkpoint does ship an MTP layer. We ran it on H100 without speculative decoding, so the +47.8% single-stream figure is our drafter recipe against NVFP4 without speculation; the c128 and c256 cells are like-for-like. And our DFlash2 drafters cost about 3× the KV cache per token, which changes our Spark guidance: for long context or non-English text, use the model's built-in MTP head.
TL;DR — DFlash2-G, our third self-trained drafter, is the first to beat the widely-used incoai reference drafter outside the noise band: 3.676 vs 3.632 mean acceptance at K=7 on the same 8× B300 (±0.015 noise). On 8× H100, our W4A16 quant runs 181.2 tok/s single-stream and 2,418 tok/s at c256. On the same rig, NVIDIA's own NVFP4 checkpoint runs 122.56 and 2,299.84, putting us +47.8% and +5.1% ahead. The first figure is against NVFP4 run without speculative decoding; see the update above. First formal vision scores: MMMU 0.747 and OCRBench 888, against 0.68 and 882 for NVFP4 on the same rig. On 2× DGX Spark, NVIDIA's NVFP4 checkpoint on its own serving stack leads all 28 cells of our head-to-head. We also correct four lines we published earlier.
The drafter: E → F → G
On 18 September we shipped DFlash2-E, our first self-trained speculative-decoding drafter. It was fully ours under Apache-2.0, but it trailed the CC-BY-NC-ND reference drafter by 2.9% on our Spark throughput primary. We said we would close the gap. Two training rounds later, all measured on the same 8× B300, same vLLM build, same 500-prompt holdout:
| Drafter | K=7, c16 | K=4, c16 | K=7, c1 | License |
|---|---|---|---|---|
| DFlash2-G (25 Sep) | 3.676 (1,410) | 3.088 (1,331) | 3.697 (315) | Apache-2.0 |
| incoai reference | 3.632 (1,425) | 3.136 (1,369) | 3.667 (328) | CC-BY-NC-ND-4.0 |
| DFlash2-F (24 Sep) | 3.626 (1,406) | 3.085 (1,339) | 3.677 (313) | Apache-2.0 |
| DFlash2-E (15 Sep) | 3.561 (1,397) | — | — | Apache-2.0 |
F reached parity with the reference: −0.006 at K=7, inside the noise. G leads it by +0.044 at K=7 and +0.030 single-stream, both outside the noise, with every draft position at or above the reference. The win is in acceptance. On raw c16 throughput on this rig the reference still edges G (1,425 vs 1,410 tok/s). A same-rig Spark A/B against the reference, where the original 2.9% gap was measured, has not been run yet. G was warm-started from F and trained for 70,000 steps on 736,675 samples. That is F's set plus 291,020 new completions of prompts we had never generated before, drawn from real user chat, tool calling, hard math, code and STEM. As before, every completion was regenerated by the target model itself. The reference drafter was never a training input; it appears here only as a measured comparison.
Two caveats. At K=4, G is still 1.5% short of the reference (3.088 vs 3.136). On 4,096-token thinking traces the two drafters are at parity (3.478 vs 3.473): both lose acceptance as the trace gets longer, and our 1K-token lead does not carry that deep. G has been the drafter-of-record on our 2× DGX Spark stack since 25 September, and it is the latency drafter in the H100 recipe below. E and F stay public with their records.
8× H100: the closeout
On 28 September we ran the canonical H100 result set on a fresh 16-GPU grid (two 8× H100 nodes). The image was our v2 build, ghcr.io/canada-quant/vllm-glm53-flash-h100:v2-w4a16-dflash2, with 8K-token prompts, 1K-token completions, and thinking ON. Output tok/s:
| Configuration | c1 | c32 | c128 | c256 |
|---|---|---|---|---|
| TP=8 + DFlash2-G K=7 (latency recipe) | 181.2 | 1,009.1 | 1,645.3 | 1,653.5 |
| TP=8, spec-decode off (aggregate recipe) | 135.4 | 1,375.9 | 2,066.8 | 2,418.1 |
| TP=8, MTP N=2 | 215.0 | 1,399.9 | 2,220.3 | 2,330.2 |
| Solo TP=4, MTP N=2 | 193.18 | 1,078.4 | 1,258.5 | 1,292.4 |
- Latency: TP=8 + DFlash2-G K=7 runs 181.2 tok/s at c1: +34% over TP=8 spec-off (135.4), and +46% over the TP=4 spec-off cell (124.1).
- Aggregate: TP=8 with spec-decode off runs 2,418 tok/s at c256; MTP N=2 is the balanced middle at 2,330.
- Best single stream on this rig: solo TP=4 with the model's own MTP head, at 193.18 tok/s. MTP shares the target's embeddings, so it avoids the drafter's standalone KV tensors.
- Don't run the DFlash2 drafter at TP=4 past c128. Its eight full-attention KV tensors shrink the KV pool about 6× (311,999 vs 1,811,949 tokens). That is a structural preemption cliff, not contention; spec-off TP=4 has no cliff.
- With any drafter on this image, cap
--max-model-lenat 131,071. 128K-token prefills fault on the speculative path; spec-off survives them.
vs NVIDIA's NVFP4, on the same H100s
NVIDIA publishes its own NVFP4 quantization of GLM-5.3-Flash. It is a Blackwell release, and Hopper has no native FP4 math, so on H100 every weight goes through FP4 emulation kernels (the engine warns about it at boot). “Not built for Hopper” is a claim worth numbers, so we ran it on the same rig with the same harness. It got two arms: our closeout flags at TP=8, the strongest NVFP4 shape this rig can produce, and NVIDIA's card recipe. The card recipe's fp8 KV cache cannot boot on Hopper, so that arm ran with bf16 KV, which is its one documented deviation.
| Runtime | c1 | c32 | c128 | c256 |
|---|---|---|---|---|
| W4A16 TP=8 + DFlash2-G K=7 (latency) | 181.2 | 1,009.1 | 1,645.3 | 1,653.5 |
| W4A16 TP=8 spec-off (aggregate) | 135.4 | 1,375.9 | 2,066.8 | 2,418.1 |
| NVIDIA NVFP4, TP=8 (our flags) | 122.56 | 1,052.71 | 2,081.13 | 2,299.84 |
| NVIDIA NVFP4, card recipe TP=4 (bf16 KV) | 81.87 | 577.28 | 845.24 | 861.95 |
We lead NVFP4's best arm by 47.8% single-stream and 5.1% at c256. c128 is a dead heat against our spec-off recipe: 2,081.13 vs 2,066.8, which is NVFP4 +0.7%. Our TP=8 MTP N=2 arm leads that cell at 2,220.3. The single-stream comparison is not like-for-like: our 181.2 uses speculative decoding, and NVFP4 ran without it because we misread its config as having no MTP layer (corrected below). The c128 and c256 cells are speculation-off on both sides. The card-recipe arm runs out of KV at TP=4 and collapses under load, with a median time to first token of 263 seconds at c256. A with-MTP NVFP4 arm on H100 has not been run yet.
First vision evals
GLM-5.3-Flash is multimodal. We keep its 24-block vision tower in BF16, and it was never touched by our text-only calibration set. Until this week we had no formal vision numbers for it at all. We ran the lmms-eval 0.7.3 task specs on the 8× H100 stack (TP=8, spec-off, greedy, seed 42), then ran NVIDIA's NVFP4 checkpoint through the same protocol on the same rig:
| Benchmark | W4A16 (ours) | NVIDIA NVFP4 |
|---|---|---|
| MMMU validation (n=900) | 0.747 (672/900) | 0.680 |
| OCRBench (n=1,000) | 888 | 882 |
The MMMU gap is +6.7 points; OCRBench is close to even. Our weakest OCRBench category is handwritten math (56/100) and our strongest is document VQA (192/200). On the text side nothing moved: AIME 2025 is 0.883 vs NVFP4's 0.900 on H100 (0.42σ, within noise), GPQA-Diamond is within noise, and GSM8K is 0.97.
Where we lost: 2× DGX Spark
On 27 September we ran our production Spark stack (W4A16 + DFlash2-G, 800K-context serve) against NVIDIA's NVFP4 checkpoint on an NVIDIA-lineage community serving stack with the incoai drafter. Both ran the same 702-request grid, the same pair of Sparks, and the same harness. Their stack led every one of the 28 cells, by 16.9% to 254.5%, with the gap widest at long context:
| Cell | W4A16 + DFlash2-G (ours) | NVIDIA NVFP4 stack |
|---|---|---|
| c1, empty context | 31.35 ± 3.52 | 34.82 ± 1.92 |
| c1 at 65K context | 15.34 | 29.19 |
| c1 at 100K context | 9.97 | 17.95 |
An isolation leg (their image with our flag values) points the gap at the weights and drafter, not the serve flags. We never ran their weights with our drafter, so the split between quant and drafter inside their lead is unmeasured. Our one measured win was stability at depth: their engine threw an HTTP 500 after the 100K sweep, while ours completed every pass twice with coherent output. We still serve our stack in production. It is ours end to end, and the drafter carries no NC/ND terms. But on Spark, NVIDIA's NVFP4 stack is faster, and we would rather tell you than let you find out.
Corrections
- “Beats NVFP4 at every tested concurrency on 4× H100” (4 Sep writeup) overstated that grid. At c8 our W4A16 was 249.94 tok/s against NVFP4's 251.14, within noise. The same-rig 8× H100 comparison above replaces it.
- “NVFP4 OOM'd all nine boots on 2× DGX Spark” is true, but it was one community checkpoint on the late-August stack, not NVIDIA's release. The “+53% single-stream / +101% at c6 vs NVFP4” Spark figures were against an earlier reference baseline, not a matched same-day run. The matched head-to-head is the Spark section above.
- “The NVFP4 checkpoint ships no MTP layer” (this post, 30 Sep) is wrong. NVIDIA's checkpoint ships one: layer 45 in BF16, with
num_nextn_predict_layers: 1in its text config. We read the wrong level of the config. The layer is missing from the checkpoint's quantization ignore list, so vLLM needs a loader fix to use it as a drafter (eugr/spark-vllm-docker PR #421 has one). Our H100 runs left speculative decoding off for NVFP4, so the +47.8% single-stream lead is our drafter recipe against NVFP4 without speculation. The c128 and c256 cells are like-for-like. Corrected 4 Oct 2026. - Our DFlash2 drafters cost about 3× the KV cache per token. DFlash2-E, -F and -G have eight full-attention layers, so their cache grows with the whole context; the reference drafter keeps a 2,048-token window. Measured on our Sparks: 366,749 tokens in an 8 GiB KV pool and 888,729 in 16 GiB with our drafter, against 1,360,420 in 9 GiB with the reference. The 1M-context Spark serve was measured with the reference drafter. With ours, 9 GiB holds about 450K tokens; we run G at 262K, and up to 800K with a 16 GiB pool. Corrected 4 Oct 2026.
Both older posts now carry a dated note pointing here, and the model card has carried the corrected framing since 27–28 September. We are making the two 4 Oct corrections on the model card and drafter cards too.
What's next
- RTX PRO 6000 (SM120) head-to-head against NVFP4 on 8× RTX PRO 6000, where NVFP4 runs natively. That is the fair Blackwell fight.
- The Spark depth gap: long-context decode is where the NVFP4 stack pulls away. For long context or non-English text, use the model's built-in MTP head rather than DFlash2-G. A community user measured 31.6 / 25.4 / 31.0 tok/s at 0 / 65K / 100K context with our weights and MTP on eugr's b12x build (different harness, GPU clock capped at 1700 MHz; not measured by us). The same user found DFlash2-G 20–30% slower on non-English text and faster on English code. Our own same-pair comparison on 2× Spark, including our image with MTP, has been running since 4 Oct.
- K=4 and deep traces: the two places DFlash2-G is still level with or behind the reference.