GLM-5.3-Flash, a month in: our drafter passes the reference, NVIDIA's NVFP4 on H100, and first vision evals

Four weeks after we shipped the W4A16 quant, we have trained two more drafters, run a full 8× H100 closeout, put NVIDIA's own NVFP4 checkpoint on the same rig, and measured the model's vision tower for the first time. Most of it went our way. One head-to-head did not, and it is in here too.

Update · 4 Oct 2026 — Two more corrections, both in the corrections section below. NVIDIA's NVFP4 checkpoint does ship an MTP layer. We ran it on H100 without speculative decoding, so the +47.8% single-stream figure is our drafter recipe against NVFP4 without speculation; the c128 and c256 cells are like-for-like. And our DFlash2 drafters cost about 3× the KV cache per token, which changes our Spark guidance: for long context or non-English text, use the model's built-in MTP head.

TL;DR — DFlash2-G, our third self-trained drafter, is the first to beat the widely-used incoai reference drafter outside the noise band: 3.676 vs 3.632 mean acceptance at K=7 on the same 8× B300 (±0.015 noise). On 8× H100, our W4A16 quant runs 181.2 tok/s single-stream and 2,418 tok/s at c256. On the same rig, NVIDIA's own NVFP4 checkpoint runs 122.56 and 2,299.84, putting us +47.8% and +5.1% ahead. The first figure is against NVFP4 run without speculative decoding; see the update above. First formal vision scores: MMMU 0.747 and OCRBench 888, against 0.68 and 882 for NVFP4 on the same rig. On 2× DGX Spark, NVIDIA's NVFP4 checkpoint on its own serving stack leads all 28 cells of our head-to-head. We also correct four lines we published earlier.

The drafter: E → F → G

On 18 September we shipped DFlash2-E, our first self-trained speculative-decoding drafter. It was fully ours under Apache-2.0, but it trailed the CC-BY-NC-ND reference drafter by 2.9% on our Spark throughput primary. We said we would close the gap. Two training rounds later, all measured on the same 8× B300, same vLLM build, same 500-prompt holdout:

Mean acceptance length (output tok/s), 500 never-trained-on prompts, thinking ON, T=1.0, TP=4, 8× B300. Run-to-run noise ±0.015.
DrafterK=7, c16K=4, c16K=7, c1License
DFlash2-G (25 Sep)3.676 (1,410)3.088 (1,331)3.697 (315)Apache-2.0
incoai reference3.632 (1,425)3.136 (1,369)3.667 (328)CC-BY-NC-ND-4.0
DFlash2-F (24 Sep)3.626 (1,406)3.085 (1,339)3.677 (313)Apache-2.0
DFlash2-E (15 Sep)3.561 (1,397)——Apache-2.0

F reached parity with the reference: −0.006 at K=7, inside the noise. G leads it by +0.044 at K=7 and +0.030 single-stream, both outside the noise, with every draft position at or above the reference. The win is in acceptance. On raw c16 throughput on this rig the reference still edges G (1,425 vs 1,410 tok/s). A same-rig Spark A/B against the reference, where the original 2.9% gap was measured, has not been run yet. G was warm-started from F and trained for 70,000 steps on 736,675 samples. That is F's set plus 291,020 new completions of prompts we had never generated before, drawn from real user chat, tool calling, hard math, code and STEM. As before, every completion was regenerated by the target model itself. The reference drafter was never a training input; it appears here only as a measured comparison.

Two caveats. At K=4, G is still 1.5% short of the reference (3.088 vs 3.136). On 4,096-token thinking traces the two drafters are at parity (3.478 vs 3.473): both lose acceptance as the trace gets longer, and our 1K-token lead does not carry that deep. G has been the drafter-of-record on our 2× DGX Spark stack since 25 September, and it is the latency drafter in the H100 recipe below. E and F stay public with their records.

8× H100: the closeout

On 28 September we ran the canonical H100 result set on a fresh 16-GPU grid (two 8× H100 nodes). The image was our v2 build, ghcr.io/canada-quant/vllm-glm53-flash-h100:v2-w4a16-dflash2, with 8K-token prompts, 1K-token completions, and thinking ON. Output tok/s:

8× H100 (SM90), isl/osl 8192/1024, thinking ON. Solo TP=4 = one TP=4 serve alone on an uncontended node.
Configurationc1c32c128c256
TP=8 + DFlash2-G K=7 (latency recipe)181.21,009.11,645.31,653.5
TP=8, spec-decode off (aggregate recipe)135.41,375.92,066.82,418.1
TP=8, MTP N=2215.01,399.92,220.32,330.2
Solo TP=4, MTP N=2193.181,078.41,258.51,292.4

vs NVIDIA's NVFP4, on the same H100s

NVIDIA publishes its own NVFP4 quantization of GLM-5.3-Flash. It is a Blackwell release, and Hopper has no native FP4 math, so on H100 every weight goes through FP4 emulation kernels (the engine warns about it at boot). “Not built for Hopper” is a claim worth numbers, so we ran it on the same rig with the same harness. It got two arms: our closeout flags at TP=8, the strongest NVFP4 shape this rig can produce, and NVIDIA's card recipe. The card recipe's fp8 KV cache cannot boot on Hopper, so that arm ran with bf16 KV, which is its one documented deviation.

Same 8× H100 rig and harness, 8192/1024, thinking ON, output tok/s, 28 Sep 2026.
Runtimec1c32c128c256
W4A16 TP=8 + DFlash2-G K=7 (latency)181.21,009.11,645.31,653.5
W4A16 TP=8 spec-off (aggregate)135.41,375.92,066.82,418.1
NVIDIA NVFP4, TP=8 (our flags)122.561,052.712,081.132,299.84
NVIDIA NVFP4, card recipe TP=4 (bf16 KV)81.87577.28845.24861.95

We lead NVFP4's best arm by 47.8% single-stream and 5.1% at c256. c128 is a dead heat against our spec-off recipe: 2,081.13 vs 2,066.8, which is NVFP4 +0.7%. Our TP=8 MTP N=2 arm leads that cell at 2,220.3. The single-stream comparison is not like-for-like: our 181.2 uses speculative decoding, and NVFP4 ran without it because we misread its config as having no MTP layer (corrected below). The c128 and c256 cells are speculation-off on both sides. The card-recipe arm runs out of KV at TP=4 and collapses under load, with a median time to first token of 263 seconds at c256. A with-MTP NVFP4 arm on H100 has not been run yet.

First vision evals

GLM-5.3-Flash is multimodal. We keep its 24-block vision tower in BF16, and it was never touched by our text-only calibration set. Until this week we had no formal vision numbers for it at all. We ran the lmms-eval 0.7.3 task specs on the 8× H100 stack (TP=8, spec-off, greedy, seed 42), then ran NVIDIA's NVFP4 checkpoint through the same protocol on the same rig:

8× H100, TP=8, spec-off, greedy, lmms-eval 0.7.3 task specs, 28 Sep 2026.
BenchmarkW4A16 (ours)NVIDIA NVFP4
MMMU validation (n=900)0.747 (672/900)0.680
OCRBench (n=1,000)888882

The MMMU gap is +6.7 points; OCRBench is close to even. Our weakest OCRBench category is handwritten math (56/100) and our strongest is document VQA (192/200). On the text side nothing moved: AIME 2025 is 0.883 vs NVFP4's 0.900 on H100 (0.42σ, within noise), GPQA-Diamond is within noise, and GSM8K is 0.97.

Where we lost: 2× DGX Spark

On 27 September we ran our production Spark stack (W4A16 + DFlash2-G, 800K-context serve) against NVIDIA's NVFP4 checkpoint on an NVIDIA-lineage community serving stack with the incoai drafter. Both ran the same 702-request grid, the same pair of Sparks, and the same harness. Their stack led every one of the 28 cells, by 16.9% to 254.5%, with the gap widest at long context:

2× DGX Spark (SM121), llama-benchy pp2048/tg128, output tok/s per stream, 27 Sep 2026.
CellW4A16 + DFlash2-G (ours)NVIDIA NVFP4 stack
c1, empty context31.35 ± 3.5234.82 ± 1.92
c1 at 65K context15.3429.19
c1 at 100K context9.9717.95

An isolation leg (their image with our flag values) points the gap at the weights and drafter, not the serve flags. We never ran their weights with our drafter, so the split between quant and drafter inside their lead is unmeasured. Our one measured win was stability at depth: their engine threw an HTTP 500 after the 100K sweep, while ours completed every pass twice with coherent output. We still serve our stack in production. It is ours end to end, and the drafter carries no NC/ND terms. But on Spark, NVIDIA's NVFP4 stack is faster, and we would rather tell you than let you find out.

Corrections

Both older posts now carry a dated note pointing here, and the model card has carried the corrected framing since 27–28 September. We are making the two 4 Oct corrections on the model card and drafter cards too.

What's next

Get it

canada-quant/GLM-5.3-Flash-W4A16-MTP (MIT) and canada-quant/GLM-5.3-Flash-DFlash2-G (Apache-2.0). The model card's BENCHMARKS.md holds every grid behind this post. 2× DGX Spark serving stack: canada-quant/vllm-glm53-flash-sm121 (Apache-2.0). Results current as of 2026-09-30; corrections added 2026-10-04.