~/blog/qwen38-flash-next-two-sparks-checkpoint

DGX Spark · part 49

[DGX Spark] Qwen3.8-Flash-Next on two DGX Sparks: 51.9 to 72.0 tok/s, and a draft vocabulary built from my own text

❯ cat --toc

TL;DR

Qwen3.8-Flash-Next NVFP4 on two DGX Sparks (TP=2) went from 51.9 to 72.0 tok/s on the 40 prompts from Part 46; the harness score rose from 0.843 to 0.876. The biggest gain came from adding my own frequently used Chinese to the MTP draft vocabulary. The stock list was built from English and code, so 82% of the Chinese tokens I sampled could never be proposed. After adding 5,606 tokens from my own text, a prompt on a topic I work on daily got 60% faster (acceptance 0.11 → 0.98 tokens per step) and everyday topics outside that corpus got 23–27% faster, from a one-line mount change. Turn-two TTFT on a 9.5k prefix fell from 3.30 s to 0.23 s. Caveat: the KV pool is 69% smaller, from a memory setting lowered after an OOM.

Preface

If you weigh yourself, you use the same scale every time, first thing in the morning, before coffee. A new scale at a different hour tells you nothing about whether you changed.

This is Part 49 of the DGX Spark series and a checkpoint against Part 46, published 2026-09-12. Sixteen days later I re-ran the same harness and the same probe on the same two machines. The first half of this post is what changed and how to reproduce it. The second half is debugging notes.

51.9 → 72.0 tok/s on the same 40 prompts, score up from 0.843 to 0.876

The setup: Qwen3.8-Flash-Next NVFP4 on two NVIDIA DGX Spark (GB10) boxes, vLLM with tensor parallel TP=2 plus expert parallel, using the MiaAI dual-Spark derivative vLLM image. Every number is a median of 3 runs.

The first row is the headline; the probe rows break the gain down by content type.

MetricPart 46 (Sep 12)Now (Sep 28)Change
40-prompt harness, median decode51.9 tok/s72.0 tok/s+38.7%
Harness auto-score0.8430.876+0.033
Probe: code61.79 tok/s86.82 tok/s+40.5%
Probe: English prose37.87 tok/s44.99 tok/s+18.8%
Probe: Traditional Chinese prose33.16 tok/s45.89 tok/s+38.4%

The individual harness runs were 56.6 / 51.7 / 51.9 in Part 46 and 77.9 / 71.2 / 72.0 now. Both sets varied by about 9% run to run. All 8 harness categories got faster: coding went from 58.1 to 82.9, and prose, the slowest category, from 38.0 to 48.6.

The auto-score comes from a grader that executes the generated Python for coding prompts. The prose category score moved from 0.4 to 0.6 and the rest stayed mostly flat (per-run: 0.855 / 0.880 / 0.876). This is not a full quality evaluation. It does show that nothing got worse.

Five changes in 16 days, and one of them made Chinese slower

Five changes landed between the two measurements:

  1. Sep 22: switched to the MiaAI dual-Spark vLLM build. The old stack had zero prefix-cache hits across requests.
  2. Sep 22: adaptive MTP with 4 draft slots and a 47,149-token draft vocabulary. MiaAI had measured a speedup with it, and code did benefit.
  3. Sep 24: gpu-memory-utilization from 0.835 to 0.68. An earlyoom cascade took the service down for 17 minutes. At 0.78 the head node's available memory still dropped to 1.8 GiB.
  4. Sep 28: prefix-cache alignment fix, plus GDN weights read as FP8.
  5. Sep 28: Chinese tokens added to the draft vocabulary.

The intermediate rounds used a different probe at different times, so they only show a trend. For Chinese the trend was down: the Chinese probe fell from 33.16 tok/s at Part 46 to 27.6 on the evening of Sep 27, and nobody noticed. The cause was change 2, the Sep 22 draft vocabulary.

Chinese was slow because 82% of its tokens were not on the draft list

MTP (multi-token prediction: a small head built into the model guesses the next few tokens each step) works like this. The draft head proposes up to 4 tokens. The target model reads its weights once, verifies every proposed position in that single pass, and keeps the guesses that match. Machines that are bound by memory bandwidth gain the most, because a weight read costs the same whether it verifies one position or five.

The MiaAI recipe makes each draft step cheaper by shrinking the draft head's proposal set. The tokenizer has 248,077 ids; the recipe limits proposals to a 47,149-token list in draft_vocab_en_code_47k.txt. As the name says, that list was built from English and code.

I measured accepted draft tokens per step by language:

ContentAccepted tokens per stepFirst-slot hit rate
Code3.15497.4%
English1.22659.4%
Chinese0.11411.4%

On four Chinese paragraphs (Traditional and Simplified: a technical explanation, operating steps, a product reply, a tutorial), 82.26% of the tokens were not on the list. The draft head cannot propose a token it is not allowed to see, so for Chinese the target model was back to emitting about one token per step, while still paying for the draft.

Diagram: the target model reads its weights once per step while the draft head guesses 4 tokens from a 47,149-token English and code list. 82% of Chinese tokens are not on that list and get rejected. Adding Chinese tokens grows the list to 52,755 ids.
With Chinese tokens added to the draft list (47,149 to 52,755 ids), accepted tokens per step on Chinese went from 0.11 to 0.98.

Build the draft vocabulary from your own text: two commands and one bind mount

The fix uses a tool that already ships in the MiaAI repo: files/build_draft_vocab.py. It reads a corpus, counts tokens, and writes an id list. I built a Chinese list and took the union with the existing one, so no English or code token is lost.

The corpus is the point: use text you actually work with. Mine was Chinese documentation already on the machine, three model-written Chinese passages on topics I routinely ask about, and Simplified Chinese versions of all of it: 174,490 tokens. If your server keeps logs, real prompts and responses will match your workload even better. Run the script inside the serving container so it uses the same tokenizer the server does.

docker cp zh_corpus_both.txt vllm_qwen38fn_miaai:/tmp/
docker cp build_draft_vocab.py vllm_qwen38fn_miaai:/tmp/
docker exec vllm_qwen38fn_miaai python3 /tmp/build_draft_vocab.py /tmp/zh_corpus_both.txt --out /tmp/zh_candidates_both.txt --size 65536

Then merge the new candidates into the existing list:

docker exec vllm_qwen38fn_miaai python3 -c "a={int(x) for x in open('/etc/vllm-draft-vocab.txt')}; b={int(x) for x in open('/tmp/zh_candidates_both.txt')}; open('/tmp/draft_vocab_en_code_zh_both.txt','w').writelines(f'{x}\n' for x in sorted(a|b)); print(len(a), len(b), len(b-a), len(a|b))"

Output:

47149 9749 5606 52755

The old list had 47,149 ids and the Chinese candidates had 9,749, of which 5,606 were new. The merged list has 52,755 ids.

The launcher change is one bind-mount line. The two environment variables were already there:

-v /home/coolthor/qwen38fn-tuning/hunt-20260927/zhvocab-compile/draft_vocab_en_code_zh_both.txt:/etc/vllm-draft-vocab.txt:ro
-e VLLM_MTP_DRAFT_VOCAB=/etc/vllm-draft-vocab.txt
-e VLLM_MTP_DRAFT_VOCAB_BALANCE=1

VLLM_MTP_DRAFT_VOCAB_BALANCE=1 splits the list evenly across the two TP ranks. It comes from MiaAI's unmerged PR #45, which is the overlay I run. The main branch does not have it, so drop that line if you are on main. An odd total is fine: both hosts log balanced=True.

Output doesn't change. Per the tool's own docs, a token missing from the list just gets rejected as a draft; it never becomes a wrong token. I checked three prompts at temperature 0 and the output was byte-identical before and after.

Chinese decode: +60% on a prompt from my usual topics, +23–27% on everyday ones

This was the biggest gain of the round, from one bind-mount line, with byte-identical output on three test prompts. How much faster depends on how close the prompt is to the corpus. All at temperature 0, 400 tokens, median of 3, old list versus new list:

PromptOld listNew listChange
Chinese probe (bandwidth and tensor parallelism)27.64 tok/s44.31 tok/s+60%
Traditional Chinese: home-style chicken recipe30.0 tok/s36.9 tok/s+23%
Traditional Chinese: planning a mountain sunrise trip29.8 tok/s37.3 tok/s+25%
Simplified Chinese: choosing a phone for an elderly parent31.7 tok/s40.3 tok/s+27%

On the probe, acceptance went from 0.114 to 0.978 tokens per step. On the three new prompts it went from 0.13–0.22 to 0.36–0.64.

The probe asks about bandwidth and tensor parallelism, which is what I work on every day and the same kind of text the list was built from. The other three are everyday topics that are not in the corpus; I wrote them after building the list, and they still rose 23–27% with acceptance up 2.6–3x. Across these four prompts, the closer the list is to what you actually run, the bigger the gain.

English and code held steady. English acceptance went from 1.226 to 1.103 per step, but English speed was flat (44.74 to 44.98 tok/s). Code went from 83.27 to 82.32, within run-to-run spread.

Turn two of a conversation: TTFT from 3.30 s to 0.23 s

Coding agents resend the whole conversation on every turn, so turn two should be almost free if the prefix cache works. On this build it was not. I measured turn two as the same prefix plus one follow-up:

Prefix lengthBefore: turn-2 TTFT / tokens hitAfter: turn-2 TTFT / tokens hit
9.5k tokens3.30 s / 00.23 s / 9,464
30k tokens4.73 s / 16,9600.25 s / 29,992

Before the fix, turn two of the 9.5k conversation was as slow as turn one.

The cause is the hybrid architecture. Some of the layers are GDN, a Mamba-style recurrent layer. A recurrent layer can only resume from a position where its state was saved. The cache block is 848 tokens, and the state is saved only when a prefill chunk stops exactly on a block boundary. The scheduler aligned chunk stops to a configured value of 8, so the last segment of the prompt never ended on a reusable boundary, and turn two recomputed it.

Diagram: cache blocks are 848 tokens. With chunk stops aligned every 8 tokens, the last segment's recurrent state is never saved and turn two recomputes it. With stops aligned to 848, turn two hits the cache.
Aligning prefill chunk stops to the 848-token cache block took turn-two TTFT on a 9.5k-token prefix from 3.30 s to 0.23 s.

The fix has two parts:

  1. Align chunk stops to 848.
  2. Borrow the wider tail-lookup margin from upstream vLLM PR #57180 (unmerged at the time of writing), and subtract the 3 tokens that MTP re-fills when the tail is registered. The subtraction is my own finding; it is not in the PR.

The change touches the scheduler and two cache-manager files, mounted into the container as overlays. There is no config flag for it.

The costs: a 69% smaller KV pool and FP8 GDN weights

KV pool. It went from 2,128,927 to 665,924 tokens (−69%). The speedups did not cause this. The Sep 24 change of gpu-memory-utilization from 0.835 to 0.68 did. On a DGX Spark the CPU and GPU share memory, so that headroom is shared with the host. At 262K tokens per request the pool holds 2.54 requests. The request queue stayed at zero.

FP8 GDN weights. The GDN weights are converted to FP8 (E4M3) at runtime. They are read on every step, so halving the bytes read made code, Chinese and English each 5–7% faster. This does change the target model's numerics. Code output was byte-identical. Chinese and English responses were worded differently, with the same meaning. Chinese style, tool calling and a 28k-token needle retrieval all passed, and the harness score held. If you need byte-stable output, this change can be removed on its own without touching the others.

Deep dive: debugging notes

You can skip this section; everything above works without it. It is here for people reproducing the setup and for AI assistants reading the page on their behalf.

Decode runs at 80–87% of memory bandwidth

Conclusion: decode is close to the bandwidth wall, so reducing bytes read is the lever.

How I measured: PyTorch profiler on an isolated code prompt. One MTP step takes 51 ms. The GPU is busy 95–96% of the time: dense GEMM 60–62%, MoE GEMM 22%, NCCL 8–10%.

Effective bandwidth is weight bytes divided by kernel time, using per-GPU weight sizes at TP=2. The GB10 spec is 273 GB/s.

KernelShape (BF16)BytesTimeEffective bandwidth
GDN QKVZ8192×256041.9 MB191 μs219 GB/s (80%)
lm_head124160×2560635.7 MB2,677 μs237 GB/s (87%)

That is why GDN went to FP8: it reads half the bytes. An E4M3 FP8 microbenchmark at M=5, including input quantization, went from 179 to 97 μs.

+15% GPU clock, no change in decode

The worker had been capped at 2,200 MHz because of a GB10 defect that hard-powers-off the box under sustained load (Part 48). I removed the cap for this test. Clock under load went from 2,177 to 2,502 MHz (+15%). Code decode over 5 runs: 81.3 versus 82.0 tok/s.

Takeaway: the workload is bandwidth-bound, not compute-bound; the cap costs nothing on decode.

Custom skinny GEMMs saved less than 1% of a step

MTP verification multiplies only 1–5 rows at a time, so I tried custom kernels for that shape:

GEMMStockCustom
GDN176.4 μs165.8 μs (−6.0%)
HC28.3 μs27.8 μs
Output71.5 μs65.0 μs

Takeaway: combined, these are under 1% of a 51 ms step. None shipped.

Conclusion: prefill is limited by host-side code and all-reduce latency, not by link bandwidth.

How I measured: a 12.5k-token prefill. The head node's forward pass takes 5.578 s; GPU kernels account for 3.278 s of it. Gaps longer than 10 ms:

  • QSA indexer host code: about 0.99 s
  • aten::mul: 0.48 s
  • MoE shared path: 0.37 s
  • one cudaMalloc: 172 ms

NCCL in the same trace: 71 BF16 all-reduces of shape (12520, 2560) at 64.1 MB each, plus one 32 MB int8 all-reduce, about 4.58 GB total. The head node spent 683.6 ms in them, which is 6.70 GB/s, or 53.6 Gb/s: 29% of the 185 Gb/s link I measured in Part 46.

A standalone 64.1 MB all-reduce between the two machines takes 6.18 ms, about 83 Gb/s. So the link is not saturated; the time goes to latency and scheduling. NCCL_PROTO=Simple was slower at both 64 MB and 25.6 KB, so I rejected it.

Rejected changes

Conclusion: none of these helped at this configuration.

ChangeResult
max_num_batched_tokens 32,768Server does not start: at GMU 0.68 the KV pool cannot hold a 262K request
max_num_batched_tokens 24,57612.5k cold TTFT 23% slower
MiaAI #66 disable_eagle_block_dropNo-op on the Qwen path (no eagle group); cache hits unchanged
mamba_cache_mode=allNo change; the non-aligned path falls back to block-size hash granularity
torch.compile mode 3Not run: MiaAI PR #45 records a 103-minute compile on two GB10s that never finished (RPC timeout)
IRQs pinned to big cores, chunk 8192, GDN FlashInfer prefillRejected Sep 24; FlashInfer prefill made a four-request run 8.8% slower

Measurement noise nearly blocked the vocab change

Conclusion: a fixed 3% gate misfires when the metric's own noise exceeds 3%.

The first A/B of the new vocabulary tripped my rule "no other metric regresses by more than 3%": 12.5k cold TTFT was 3.67% slower. I went back to the old configuration and ran the same measurement again: 6.35 / 4.14 / 6.10 s. Cold TTFT swings about 50% on its own. The draft vocabulary only affects which tokens MTP proposes during decode; prefill never touches it.

The gate was meant to ask "did this make anything worse". Written as "nothing else moves more than 3%", it fires every time on a metric noisier than 3%. After checking which code path the change touches and accounting for the noise, I shipped the vocabulary.

A reboot would bring back the old stack

Conclusion: this configuration does not survive a reboot yet. Not fixed.

The worker still has the old stack's qwen38fn-tp2-worker.service enabled. The one actually in use, qwen38fn-miaai-worker.service, is disabled. After a reboot the worker would come up on the old stack.

Takeaway: if you copy this setup, enable the service you actually run and disable the old one.

Methodology

Harness (same 40 prompts as Part 46, thinking off, concurrency 1, run on the GX10 because coding prompts execute code):

cd ~/src/tonyd2wild-nvfp4/single-spark-vllm-tp1 && python3 tools/bench_categories.py http://127.0.0.1:8000 qwen38-flash-next <lane> --thinking off --concurrency 1 --tag runN

Probe (temperature 0):

python3 ~/patches/decode_probe3.py qwen38-flash-next <code|prose_en|prose_zh> 0

Rules:

  • Each run started only when vllm:num_requests_running was 0.
  • Accepted tokens per step = the change in spec_decode_num_drafts_total and spec_decode_num_accepted_tokens_total across a run, kept only if request_success_total rose by exactly 1.
  • Turn-two hits = the increment of prefix_cache_hits_total.
  • 3 runs per configuration, median reported.

Also in this series: Part 46: Qwen3.8-Flash-Next TP=2 on two DGX Sparks, 51.9 tok/s · Part 45: single-node NVFP4 recipe, 41.7 tok/s

FAQ

Why is Qwen3.8-Flash-Next with MTP so much slower on Chinese than on English?
If you use the reduced draft vocabulary from the MiaAI dual-Spark recipe (draft_vocab_en_code_47k.txt), that list was built from English and code only. On four Chinese paragraphs, 82% of the tokens were not on it, so the draft head could not propose them and only 0.11 tokens were accepted per step. Adding Chinese tokens with the repo's own build_draft_vocab.py raised that to 0.98, and our Chinese test prompt went from 27.6 to 44.3 tok/s.
Does adding tokens to an MTP draft vocabulary change the model's output?
No. The draft only proposes; the target model decides what is kept. A token missing from the list becomes a rejected draft, not a wrong token. Three prompts at temperature 0 produced byte-identical output before and after the swap.
Why does vLLM miss the prefix cache on the second turn of a conversation with a hybrid Mamba/GDN model?
A hybrid model can only resume from a position where its recurrent state was saved. In our build the cache block is 848 tokens, but the scheduler aligned prefill chunk stops to a configured value of 8, so the last segment's state never landed on a reusable boundary. A 9.5k-token prefix got zero hits on turn two. Aligning stops to 848 and subtracting the 3 tokens MTP re-fills took turn-two TTFT from 3.30 s to 0.23 s. Upstream vLLM PR #57180 addresses the same class of bug (reusing Mamba prompt tails under MTP).
Does raising the GPU clock on a DGX Spark speed up LLM decode?
Not for this workload. Raising the worker's clock under load from 2,177 to 2,502 MHz (+15%) left code decode at 81.3 versus 82.0 tok/s. Decode is memory-bandwidth bound: the lm_head read already runs at 87% of the 273 GB/s spec.

Read next

Don't miss the next one

Subscribe, and you won't.

One-click unsubscribe anytime.