DGX Spark · part 49
[DGX Spark] Qwen3.8-Flash-Next on two DGX Sparks: 51.9 to 72.0 tok/s, and a draft vocabulary built from my own text
❯ cat --toc
- Preface
- 51.9 → 72.0 tok/s on the same 40 prompts, score up from 0.843 to 0.876
- Five changes in 16 days, and one of them made Chinese slower
- Chinese was slow because 82% of its tokens were not on the draft list
- Build the draft vocabulary from your own text: two commands and one bind mount
- Chinese decode: +60% on a prompt from my usual topics, +23–27% on everyday ones
- Turn two of a conversation: TTFT from 3.30 s to 0.23 s
- The costs: a 69% smaller KV pool and FP8 GDN weights
- Deep dive: debugging notes
- Decode runs at 80–87% of memory bandwidth
- +15% GPU clock, no change in decode
- Custom skinny GEMMs saved less than 1% of a step
- Prefill: host-side gaps dominate, and NCCL uses 29% of the link
- Rejected changes
- Measurement noise nearly blocked the vocab change
- A reboot would bring back the old stack
- Methodology
TL;DR
Qwen3.8-Flash-Next NVFP4 on two DGX Sparks (TP=2) went from 51.9 to 72.0 tok/s on the 40 prompts from Part 46; the harness score rose from 0.843 to 0.876. The biggest gain came from adding my own frequently used Chinese to the MTP draft vocabulary. The stock list was built from English and code, so 82% of the Chinese tokens I sampled could never be proposed. After adding 5,606 tokens from my own text, a prompt on a topic I work on daily got 60% faster (acceptance 0.11 → 0.98 tokens per step) and everyday topics outside that corpus got 23–27% faster, from a one-line mount change. Turn-two TTFT on a 9.5k prefix fell from 3.30 s to 0.23 s. Caveat: the KV pool is 69% smaller, from a memory setting lowered after an OOM.
Preface
If you weigh yourself, you use the same scale every time, first thing in the morning, before coffee. A new scale at a different hour tells you nothing about whether you changed.
This is Part 49 of the DGX Spark series and a checkpoint against Part 46, published 2026-09-12. Sixteen days later I re-ran the same harness and the same probe on the same two machines. The first half of this post is what changed and how to reproduce it. The second half is debugging notes.
51.9 → 72.0 tok/s on the same 40 prompts, score up from 0.843 to 0.876
The setup: Qwen3.8-Flash-Next NVFP4 on two NVIDIA DGX Spark (GB10) boxes, vLLM with tensor parallel TP=2 plus expert parallel, using the MiaAI dual-Spark derivative vLLM image. Every number is a median of 3 runs.
The first row is the headline; the probe rows break the gain down by content type.
| Metric | Part 46 (Sep 12) | Now (Sep 28) | Change |
|---|---|---|---|
| 40-prompt harness, median decode | 51.9 tok/s | 72.0 tok/s | +38.7% |
| Harness auto-score | 0.843 | 0.876 | +0.033 |
| Probe: code | 61.79 tok/s | 86.82 tok/s | +40.5% |
| Probe: English prose | 37.87 tok/s | 44.99 tok/s | +18.8% |
| Probe: Traditional Chinese prose | 33.16 tok/s | 45.89 tok/s | +38.4% |
The individual harness runs were 56.6 / 51.7 / 51.9 in Part 46 and 77.9 / 71.2 / 72.0 now. Both sets varied by about 9% run to run. All 8 harness categories got faster: coding went from 58.1 to 82.9, and prose, the slowest category, from 38.0 to 48.6.
The auto-score comes from a grader that executes the generated Python for coding prompts. The prose category score moved from 0.4 to 0.6 and the rest stayed mostly flat (per-run: 0.855 / 0.880 / 0.876). This is not a full quality evaluation. It does show that nothing got worse.
Five changes in 16 days, and one of them made Chinese slower
Five changes landed between the two measurements:
- Sep 22: switched to the MiaAI dual-Spark vLLM build. The old stack had zero prefix-cache hits across requests.
- Sep 22: adaptive MTP with 4 draft slots and a 47,149-token draft vocabulary. MiaAI had measured a speedup with it, and code did benefit.
- Sep 24:
gpu-memory-utilizationfrom 0.835 to 0.68. An earlyoom cascade took the service down for 17 minutes. At 0.78 the head node's available memory still dropped to 1.8 GiB. - Sep 28: prefix-cache alignment fix, plus GDN weights read as FP8.
- Sep 28: Chinese tokens added to the draft vocabulary.
The intermediate rounds used a different probe at different times, so they only show a trend. For Chinese the trend was down: the Chinese probe fell from 33.16 tok/s at Part 46 to 27.6 on the evening of Sep 27, and nobody noticed. The cause was change 2, the Sep 22 draft vocabulary.
Chinese was slow because 82% of its tokens were not on the draft list
MTP (multi-token prediction: a small head built into the model guesses the next few tokens each step) works like this. The draft head proposes up to 4 tokens. The target model reads its weights once, verifies every proposed position in that single pass, and keeps the guesses that match. Machines that are bound by memory bandwidth gain the most, because a weight read costs the same whether it verifies one position or five.
The MiaAI recipe makes each draft step cheaper by shrinking the draft head's proposal set. The tokenizer has 248,077 ids; the recipe limits proposals to a 47,149-token list in draft_vocab_en_code_47k.txt. As the name says, that list was built from English and code.
I measured accepted draft tokens per step by language:
| Content | Accepted tokens per step | First-slot hit rate |
|---|---|---|
| Code | 3.154 | 97.4% |
| English | 1.226 | 59.4% |
| Chinese | 0.114 | 11.4% |
On four Chinese paragraphs (Traditional and Simplified: a technical explanation, operating steps, a product reply, a tutorial), 82.26% of the tokens were not on the list. The draft head cannot propose a token it is not allowed to see, so for Chinese the target model was back to emitting about one token per step, while still paying for the draft.

Build the draft vocabulary from your own text: two commands and one bind mount
The fix uses a tool that already ships in the MiaAI repo: files/build_draft_vocab.py. It reads a corpus, counts tokens, and writes an id list. I built a Chinese list and took the union with the existing one, so no English or code token is lost.
The corpus is the point: use text you actually work with. Mine was Chinese documentation already on the machine, three model-written Chinese passages on topics I routinely ask about, and Simplified Chinese versions of all of it: 174,490 tokens. If your server keeps logs, real prompts and responses will match your workload even better. Run the script inside the serving container so it uses the same tokenizer the server does.
docker cp zh_corpus_both.txt vllm_qwen38fn_miaai:/tmp/
docker cp build_draft_vocab.py vllm_qwen38fn_miaai:/tmp/
docker exec vllm_qwen38fn_miaai python3 /tmp/build_draft_vocab.py /tmp/zh_corpus_both.txt --out /tmp/zh_candidates_both.txt --size 65536
Then merge the new candidates into the existing list:
docker exec vllm_qwen38fn_miaai python3 -c "a={int(x) for x in open('/etc/vllm-draft-vocab.txt')}; b={int(x) for x in open('/tmp/zh_candidates_both.txt')}; open('/tmp/draft_vocab_en_code_zh_both.txt','w').writelines(f'{x}\n' for x in sorted(a|b)); print(len(a), len(b), len(b-a), len(a|b))"
Output:
47149 9749 5606 52755
The old list had 47,149 ids and the Chinese candidates had 9,749, of which 5,606 were new. The merged list has 52,755 ids.
The launcher change is one bind-mount line. The two environment variables were already there:
-v /home/coolthor/qwen38fn-tuning/hunt-20260927/zhvocab-compile/draft_vocab_en_code_zh_both.txt:/etc/vllm-draft-vocab.txt:ro
-e VLLM_MTP_DRAFT_VOCAB=/etc/vllm-draft-vocab.txt
-e VLLM_MTP_DRAFT_VOCAB_BALANCE=1
VLLM_MTP_DRAFT_VOCAB_BALANCE=1 splits the list evenly across the two TP ranks. It comes from MiaAI's unmerged PR #45, which is the overlay I run. The main branch does not have it, so drop that line if you are on main. An odd total is fine: both hosts log balanced=True.
Output doesn't change. Per the tool's own docs, a token missing from the list just gets rejected as a draft; it never becomes a wrong token. I checked three prompts at temperature 0 and the output was byte-identical before and after.
Chinese decode: +60% on a prompt from my usual topics, +23–27% on everyday ones
This was the biggest gain of the round, from one bind-mount line, with byte-identical output on three test prompts. How much faster depends on how close the prompt is to the corpus. All at temperature 0, 400 tokens, median of 3, old list versus new list:
| Prompt | Old list | New list | Change |
|---|---|---|---|
| Chinese probe (bandwidth and tensor parallelism) | 27.64 tok/s | 44.31 tok/s | +60% |
| Traditional Chinese: home-style chicken recipe | 30.0 tok/s | 36.9 tok/s | +23% |
| Traditional Chinese: planning a mountain sunrise trip | 29.8 tok/s | 37.3 tok/s | +25% |
| Simplified Chinese: choosing a phone for an elderly parent | 31.7 tok/s | 40.3 tok/s | +27% |
On the probe, acceptance went from 0.114 to 0.978 tokens per step. On the three new prompts it went from 0.13–0.22 to 0.36–0.64.
The probe asks about bandwidth and tensor parallelism, which is what I work on every day and the same kind of text the list was built from. The other three are everyday topics that are not in the corpus; I wrote them after building the list, and they still rose 23–27% with acceptance up 2.6–3x. Across these four prompts, the closer the list is to what you actually run, the bigger the gain.
English and code held steady. English acceptance went from 1.226 to 1.103 per step, but English speed was flat (44.74 to 44.98 tok/s). Code went from 83.27 to 82.32, within run-to-run spread.
Turn two of a conversation: TTFT from 3.30 s to 0.23 s
Coding agents resend the whole conversation on every turn, so turn two should be almost free if the prefix cache works. On this build it was not. I measured turn two as the same prefix plus one follow-up:
| Prefix length | Before: turn-2 TTFT / tokens hit | After: turn-2 TTFT / tokens hit |
|---|---|---|
| 9.5k tokens | 3.30 s / 0 | 0.23 s / 9,464 |
| 30k tokens | 4.73 s / 16,960 | 0.25 s / 29,992 |
Before the fix, turn two of the 9.5k conversation was as slow as turn one.
The cause is the hybrid architecture. Some of the layers are GDN, a Mamba-style recurrent layer. A recurrent layer can only resume from a position where its state was saved. The cache block is 848 tokens, and the state is saved only when a prefill chunk stops exactly on a block boundary. The scheduler aligned chunk stops to a configured value of 8, so the last segment of the prompt never ended on a reusable boundary, and turn two recomputed it.

The fix has two parts:
- Align chunk stops to 848.
- Borrow the wider tail-lookup margin from upstream vLLM PR #57180 (unmerged at the time of writing), and subtract the 3 tokens that MTP re-fills when the tail is registered. The subtraction is my own finding; it is not in the PR.
The change touches the scheduler and two cache-manager files, mounted into the container as overlays. There is no config flag for it.
The costs: a 69% smaller KV pool and FP8 GDN weights
KV pool. It went from 2,128,927 to 665,924 tokens (−69%). The speedups did not cause this. The Sep 24 change of gpu-memory-utilization from 0.835 to 0.68 did. On a DGX Spark the CPU and GPU share memory, so that headroom is shared with the host. At 262K tokens per request the pool holds 2.54 requests. The request queue stayed at zero.
FP8 GDN weights. The GDN weights are converted to FP8 (E4M3) at runtime. They are read on every step, so halving the bytes read made code, Chinese and English each 5–7% faster. This does change the target model's numerics. Code output was byte-identical. Chinese and English responses were worded differently, with the same meaning. Chinese style, tool calling and a 28k-token needle retrieval all passed, and the harness score held. If you need byte-stable output, this change can be removed on its own without touching the others.
Deep dive: debugging notes
You can skip this section; everything above works without it. It is here for people reproducing the setup and for AI assistants reading the page on their behalf.
Decode runs at 80–87% of memory bandwidth
Conclusion: decode is close to the bandwidth wall, so reducing bytes read is the lever.
How I measured: PyTorch profiler on an isolated code prompt. One MTP step takes 51 ms. The GPU is busy 95–96% of the time: dense GEMM 60–62%, MoE GEMM 22%, NCCL 8–10%.
Effective bandwidth is weight bytes divided by kernel time, using per-GPU weight sizes at TP=2. The GB10 spec is 273 GB/s.
| Kernel | Shape (BF16) | Bytes | Time | Effective bandwidth |
|---|---|---|---|---|
| GDN QKVZ | 8192×2560 | 41.9 MB | 191 μs | 219 GB/s (80%) |
| lm_head | 124160×2560 | 635.7 MB | 2,677 μs | 237 GB/s (87%) |
That is why GDN went to FP8: it reads half the bytes. An E4M3 FP8 microbenchmark at M=5, including input quantization, went from 179 to 97 μs.
+15% GPU clock, no change in decode
The worker had been capped at 2,200 MHz because of a GB10 defect that hard-powers-off the box under sustained load (Part 48). I removed the cap for this test. Clock under load went from 2,177 to 2,502 MHz (+15%). Code decode over 5 runs: 81.3 versus 82.0 tok/s.
Takeaway: the workload is bandwidth-bound, not compute-bound; the cap costs nothing on decode.
Custom skinny GEMMs saved less than 1% of a step
MTP verification multiplies only 1–5 rows at a time, so I tried custom kernels for that shape:
| GEMM | Stock | Custom |
|---|---|---|
| GDN | 176.4 μs | 165.8 μs (−6.0%) |
| HC | 28.3 μs | 27.8 μs |
| Output | 71.5 μs | 65.0 μs |
Takeaway: combined, these are under 1% of a 51 ms step. None shipped.
Prefill: host-side gaps dominate, and NCCL uses 29% of the link
Conclusion: prefill is limited by host-side code and all-reduce latency, not by link bandwidth.
How I measured: a 12.5k-token prefill. The head node's forward pass takes 5.578 s; GPU kernels account for 3.278 s of it. Gaps longer than 10 ms:
- QSA indexer host code: about 0.99 s
aten::mul: 0.48 s- MoE shared path: 0.37 s
- one
cudaMalloc: 172 ms
NCCL in the same trace: 71 BF16 all-reduces of shape (12520, 2560) at 64.1 MB each, plus one 32 MB int8 all-reduce, about 4.58 GB total. The head node spent 683.6 ms in them, which is 6.70 GB/s, or 53.6 Gb/s: 29% of the 185 Gb/s link I measured in Part 46.
A standalone 64.1 MB all-reduce between the two machines takes 6.18 ms, about 83 Gb/s. So the link is not saturated; the time goes to latency and scheduling. NCCL_PROTO=Simple was slower at both 64 MB and 25.6 KB, so I rejected it.
Rejected changes
Conclusion: none of these helped at this configuration.
| Change | Result |
|---|---|
max_num_batched_tokens 32,768 | Server does not start: at GMU 0.68 the KV pool cannot hold a 262K request |
max_num_batched_tokens 24,576 | 12.5k cold TTFT 23% slower |
MiaAI #66 disable_eagle_block_drop | No-op on the Qwen path (no eagle group); cache hits unchanged |
mamba_cache_mode=all | No change; the non-aligned path falls back to block-size hash granularity |
| torch.compile mode 3 | Not run: MiaAI PR #45 records a 103-minute compile on two GB10s that never finished (RPC timeout) |
| IRQs pinned to big cores, chunk 8192, GDN FlashInfer prefill | Rejected Sep 24; FlashInfer prefill made a four-request run 8.8% slower |
Measurement noise nearly blocked the vocab change
Conclusion: a fixed 3% gate misfires when the metric's own noise exceeds 3%.
The first A/B of the new vocabulary tripped my rule "no other metric regresses by more than 3%": 12.5k cold TTFT was 3.67% slower. I went back to the old configuration and ran the same measurement again: 6.35 / 4.14 / 6.10 s. Cold TTFT swings about 50% on its own. The draft vocabulary only affects which tokens MTP proposes during decode; prefill never touches it.
The gate was meant to ask "did this make anything worse". Written as "nothing else moves more than 3%", it fires every time on a metric noisier than 3%. After checking which code path the change touches and accounting for the noise, I shipped the vocabulary.
A reboot would bring back the old stack
Conclusion: this configuration does not survive a reboot yet. Not fixed.
The worker still has the old stack's qwen38fn-tp2-worker.service enabled. The one actually in use, qwen38fn-miaai-worker.service, is disabled. After a reboot the worker would come up on the old stack.
Takeaway: if you copy this setup, enable the service you actually run and disable the old one.
Methodology
Harness (same 40 prompts as Part 46, thinking off, concurrency 1, run on the GX10 because coding prompts execute code):
cd ~/src/tonyd2wild-nvfp4/single-spark-vllm-tp1 && python3 tools/bench_categories.py http://127.0.0.1:8000 qwen38-flash-next <lane> --thinking off --concurrency 1 --tag runN
Probe (temperature 0):
python3 ~/patches/decode_probe3.py qwen38-flash-next <code|prose_en|prose_zh> 0
Rules:
- Each run started only when
vllm:num_requests_runningwas 0. - Accepted tokens per step = the change in
spec_decode_num_drafts_totalandspec_decode_num_accepted_tokens_totalacross a run, kept only ifrequest_success_totalrose by exactly 1. - Turn-two hits = the increment of
prefix_cache_hits_total. - 3 runs per configuration, median reported.
Also in this series: Part 46: Qwen3.8-Flash-Next TP=2 on two DGX Sparks, 51.9 tok/s · Part 45: single-node NVFP4 recipe, 41.7 tok/s
FAQ
- Why is Qwen3.8-Flash-Next with MTP so much slower on Chinese than on English?
- If you use the reduced draft vocabulary from the MiaAI dual-Spark recipe (draft_vocab_en_code_47k.txt), that list was built from English and code only. On four Chinese paragraphs, 82% of the tokens were not on it, so the draft head could not propose them and only 0.11 tokens were accepted per step. Adding Chinese tokens with the repo's own build_draft_vocab.py raised that to 0.98, and our Chinese test prompt went from 27.6 to 44.3 tok/s.
- Does adding tokens to an MTP draft vocabulary change the model's output?
- No. The draft only proposes; the target model decides what is kept. A token missing from the list becomes a rejected draft, not a wrong token. Three prompts at temperature 0 produced byte-identical output before and after the swap.
- Why does vLLM miss the prefix cache on the second turn of a conversation with a hybrid Mamba/GDN model?
- A hybrid model can only resume from a position where its recurrent state was saved. In our build the cache block is 848 tokens, but the scheduler aligned prefill chunk stops to a configured value of 8, so the last segment's state never landed on a reusable boundary. A 9.5k-token prefix got zero hits on turn two. Aligning stops to 848 and subtracting the 3 tokens MTP re-fills took turn-two TTFT from 3.30 s to 0.23 s. Upstream vLLM PR #57180 addresses the same class of bug (reusing Mamba prompt tails under MTP).
- Does raising the GPU clock on a DGX Spark speed up LLM decode?
- Not for this workload. Raising the worker's clock under load from 2,177 to 2,502 MHz (+15%) left code decode at 81.3 versus 82.0 tok/s. Decode is memory-bandwidth bound: the lm_head read already runs at 87% of the 273 GB/s spec.
Read next
- 2026-05-06Liftoff: Gemma 4 hits 670 tok/s aggregate on DGX Spark (108 tok/s single-stream)
Google announced Multi-Token Prediction drafters for Gemma 4 on 2026-05-05. The vLLM PR was opened and approved the same day; a preview Docker image shipped hours later. I tested it on DGX Spark: Gemma 4 26B-A4B-it FP8 + MTP γ=4 hits 108.78 tok/s single-stream (2.66× baseline), 674.28 tok/s aggregate at concurrency=8. One undocumented trap: the drafter pairs with -it, not base.
- 2026-09-12[vLLM] Qwen3.8-Flash-Next TP=2 on Two DGX Sparks: 51.9 tok/s
A second DGX Spark takes Qwen3.8-Flash-Next NVFP4 from 41.7 to 51.9 tok/s and the KV pool to 2.07x, plus the ConnectX-7 cabling steps the playbook leaves out.
- 2026-09-06[Benchmark] Qwen3.8-Flash-Next NVFP4 on a DGX Spark: 41.7 tok/s, RAM for Traffic, Disk for the Dictionary
NVIDIA's NVFP4 checkpoint at 41.7 tok/s on one DGX Spark via nine bind-mounted vLLM files, plus six figures on why the 47.68 GiB n-gram table lives on NVMe.
- 2026-08-23[Benchmark] Qwen3.8-27B hits 65 tok/s on one DGX Spark — 12.3 without speculative decoding
GB10 gives you 273 GB/s and the model reads 18.77 GB per token, so the ceiling is 14.54 tok/s. Measured 12.3 — 85% of it. Kernel swaps can't recover the rest. Speculative decoding gets 5.28x, and the win came from one flag.
Don't miss the next one
Subscribe, and you won't.
One-click unsubscribe anytime.