改裝 2080 Ti 22G · part 22
NVFP4 Without FP4 Hardware: Qwen3.8-27B on Two RTX 2080 Tis at 151.7 tok/s
❯ cat --toc
- Same books, smaller editions: the model got faster by getting lighter
- Self-made NVFP4 on two 2080 Tis: code 116.1 → 151.7 tok/s, 4 requests at once
- Why NVFP4 helps without FP4 hardware: 29 GB → 20 GB read per token
- Three NVFP4 checkpoints, more than 2× apart: 65.3, 136.0 and 149.9 tok/s on code
- Convert BF16 to NVFP4 yourself: one numpy script, no GPU, no calibration
- Launch with KV fp8 and `--max_batch 4`, same FastLLM build
- KV fp8 and 4 streams: 293 tok/s combined, single stream still ~150
- The cost: a 54K-token prompt prefills in 50.6 s instead of 40.5 s
- Capability: 131 / 128 on 140 questions, against 132 / 127 for FP8
- Deep dive: the shared global scale, bit-exact checks, the DFlash re-sweep and two 4-bit dead ends
- The shared global scale: 128 groups, and `in_proj_a` is the one that bites
- Bit-exact check against lyf's ModelOpt output
- DFlash re-sweep on lyf's checkpoint: 8 draft tokens with NVFP4 draft quantization is still best
- KV and concurrency detail: acceptance, prefill and short-request latency per setting
- orcarouter's official NVFP4: loads fine, runs NVFP4 kernels, and is still slow
- Two 4-bit dead ends with FastLLM's own INT4
- `--dtype fp8_e4m3` doesn't convert a BF16 checkpoint
- Pinned versions
- Limitations and next steps
TL;DR
I converted Qwen3.8-27B from BF16 to NVFP4 myself and serve it on two modded 2080 Ti 22G with FastLLM + DFlash2. Code went from 116.1 to 151.7 tok/s, math 135.5 → 178.8, prose 74.2 → 101.9, and the server now takes 4 requests at once (297 tok/s combined). The 2080 Ti has no FP4 hardware. It still wins because generation is bound by reading weights, and NVFP4 cuts that read from 29 GB to 20 GB per token. orcarouter's official NVFP4 is slower than FP8 here. Cost: a 54K-token prompt prefills 25% slower. 140-question score: 131/128 vs 132/127.

Same books, smaller editions: the model got faster by getting lighter
Movers bill by the trip, not by how many books you own. Put the same books in paperback and each box holds more pages, so the move takes fewer trips.
Text generation works the same way. Every token the model produces reads the whole set of weights out of VRAM once. With FP8 weights, Qwen3.8-27B reads 29 GB per token. With NVFP4 it reads 20 GB.
Part 21 left this machine running a self-built FastLLM with an FP8 model, two concurrent requests, and a scheduler that lets short requests cut in. The 2080 Ti is Turing and has no FP4 hardware, so I expected NVFP4 to be pointless on it. This part switches the main model to an NVFP4 checkpoint I converted myself from BF16, moves the KV cache to fp8, and raises concurrency to 4. Code went from 116.1 to 151.7 tok/s.
Self-made NVFP4 on two 2080 Tis: code 116.1 → 151.7 tok/s, 4 requests at once
Same two cards (TP=2), the same self-built FastLLM as part 21, the same benchmark script. Method: real chat requests, temperature 0, 512 output tokens, code / math / prose prompts, 3 runs each, median. Read the first row for speed and the last for whether anything broke.
| FP8 (part 21) | self-made NVFP4 | |
|---|---|---|
| code / math / prose tok/s | 116.1 / 135.5 / 74.2 | 151.7 / 178.8 / 101.9 |
| concurrent requests | 2 | 4 (297 tok/s combined) |
| main model size | 29 GB | 20 GB |
| 54K-token prompt prefill | 40.5 s | 50.6 s |
| 140 questions (no thinking / thinking+low) | 132 / 127 | 131 / 128 |
With two requests at once, each stream went from about 88 to 107.8 tok/s. The per-request context limit went up from 222,080 to 257,920 tokens. When a short request arrives during a long prompt's prefill, its first token now takes 0.52 s instead of 1.18 s, and it finishes in 5.5 s (was 3.8 s). The 140 questions ran on the same NVFP4 model before the KV and concurrency changes.
Three changes produced this:
- Main model → NVFP4, converted myself from BF16, following a format that had already proven fast on this card.
- KV cache fp4 → fp8. The smaller weights freed enough VRAM, and fp8 KV is faster under concurrency.
--max_batch2 → 4.
Why NVFP4 helps without FP4 hardware: 29 GB → 20 GB read per token
Turing tensor cores do FP16, INT8 and INT4. They don't do FP4. FastLLM stores NVFP4 weights as they are, 4-bit in VRAM, and unpacks them to FP16 inside the kernel before the tensor cores touch them. The upstream README lists NVFP4 decode kernels tuned for SM75 (the 2080 Ti) for several matrices of width 5120, which is Qwen3.8-27B's hidden size.
So NVFP4 saves bytes read, not compute. That's the right thing to save here: generation is bottlenecked on reading weights, and the compute units sit mostly idle. Without speculative decoding, code went from 32.9 tok/s (FP8, 29 GB) to 45.1 tok/s (NVFP4, 20 GB), +37%.
DFlash2, the draft model from part 20, guesses several tokens ahead and lets the main model verify 8 positions per read of its weights. That multiplies whatever the base read speed is, so the gain carries over at a similar ratio: 116.1 → 151.7 on code.
Three NVFP4 checkpoints, more than 2× apart: 65.3, 136.0 and 149.9 tok/s on code
"NVFP4" on a model card doesn't tell you which layers are NVFP4. I tried three checkpoints under the same FastLLM and the same launch flags (KV fp4, --max_batch 2). Look at the code column and the last column together.
| checkpoint | what's NVFP4 | code / math / prose | no-spec code | 140 q |
|---|---|---|---|---|
| FP8 (part 21) | none; linear layers FP8 | 116.6 / 135.8 / 74.6 | 32.9 | 132 / 127 |
| orcarouter official NVFP4 | some MLP, rest FP8 | 65.3 / 77.9 / 43.6 | 32.4 | not tested |
| lyf Heretic-ARA NVFP4 | all linear layers | 136.0 / 185.5 / 102.1 | 45.3 | 124 / 129 |
| self-made orcarouter NVFP4 | all linear layers | 149.9 / 178.0 / 104.3 | 45.1 | 131 / 128 |
orcarouter's official NVFP4 comes from the same publisher as the uncensored FP8 I was already running, but it's mixed precision: the config lists 168 NVFP4 targets (MLP) and 232 FP8 targets. Without speculation it runs 32.4, about the same as FP8. With DFlash2 it drops to 65.3 on code, slower than the FP8 checkpoint. The draft isn't the problem: its acceptance is actually higher than with FP8. My guess, not separately verified: the per-channel FP8 weights take a slower path when the main model verifies 9 tokens at once.
lyf's checkpoint was converted with NVIDIA ModelOpt, every linear layer in NVFP4, and it's fast. But it's built on a different uncensored base (Heretic ARA), and its 140-question score without thinking drops to 124, mostly on MuSR and HumanEval+.
The self-made one takes orcarouter's BF16 original and converts it into lyf's format. It uses the same base model as my FP8 build, nearly matches lyf on speed, and scores about the same as FP8.
It's also a bit faster than lyf on code. Draft acceptance at positions 1–3 is 88.8 / 73.4 / 55.0%, against lyf's 82.8 / 64.3 / 52.4%. My guess: orcarouter's uncensoring is lighter, so the output stays closer to stock Qwen3.8, which is what the DFlash2 draft was trained on.
Convert BF16 to NVFP4 yourself: one numpy script, no GPU, no calibration
You need three things:
- The BF16 original, orcarouter/Qwen3.8-27B-Uncensored, about 56 GB. It's gated: accept the terms on Hugging Face first.
- The json and yaml files from lyf's repo, without the weights. The converter copies its tensor naming, shard layout and
quantization_config. - The converter, /tools/qwen38-nvfp4-convert.py. It only needs numpy.
hf download orcarouter/Qwen3.8-27B-Uncensored --local-dir /mnt/nvme1t/models/orcarouter-qwen38-uncensored-bf16
hf download lyf/Qwen3.8-27B-Heretic-ARA-NVFP4-MTP-VL --include "*.json" "recipe.yaml" --local-dir ./nvfp4-ref
curl -O https://ai-muninn.com/tools/qwen38-nvfp4-convert.py
python3 qwen38-nvfp4-convert.py \
/mnt/nvme1t/models/orcarouter-qwen38-uncensored-bf16 \
/mnt/nvme1t/models/orcarouter-qwen38-uncensored-nvfp4-self \
--reference ./nvfp4-ref
The output is 20 GB in three shards. What the script does:
- Every linear layer → NVFP4. Values are packed as 4-bit E2M1, with one FP8 E4M3 scale per 16 values and one FP32 global scale per tensor.
- Three groups share one global scale: attention q/k/v, the four GDN input projections (
in_proj_qkv/z/a/b), and MLP gate/up. I only found this by comparing against lyf's weights. Computing each tensor's scale separately gives the right format with mismatched values. - Kept in BF16:
lm_head,embed_tokens, the vision tower, conv1d, and MTP. - No calibration. ModelOpt also calibrates activation scales (
input_global_scale) for W4A4. FastLLM on a 2080 Ti uses FP16 activations and never reads that field, so the script writes 1.0.
To check it, I decoded lyf's NVFP4 back to BF16 and re-quantized it with the script. For one attention layer, one GDN layer and one MLP layer, the packed weights, block scales and global scales were all bit-identical to lyf's. All 2,687 tensors in the output match lyf's in name, dtype, shape and shard.
Launch with KV fp8 and --max_batch 4, same FastLLM build
No rebuild. This is the production launch:
export LD_PRELOAD=/home/coolthor/nccl-install/lib/libnccl.so.2
export CUDA_VISIBLE_DEVICES=0,1
export FASTLLM_CUDA_GRAPH=0
export FASTLLM_DRAFT_QUANT=nvfp4
export FASTLLM_COOPERATIVE_LONG_PREFILL=1
~/venvs/ftllm-dual/bin/ftllm server \
-p /mnt/nvme1t/models/orcarouter-qwen38-uncensored-nvfp4-self --model_name qwen38-27b-ud \
--tp 2 --max_batch 4 --chunked_prefill_size 4096 \
--gpu_mem_ratio 0.95 --kv_cache_dtype fp8_e4m3 \
--speculative_algorithm dflash \
--speculative_draft_model_path /mnt/nvme1t/models/dflash2-qwen38-zlab \
--speculative_num_draft_tokens 8 \
--prefix_cache true \
--chat_template /mnt/nvme1t/models/qwen-fastllm-low/chat_template.jinja \
--host 0.0.0.0 --port 8082
Leave out --tokens; with it, FastLLM clamps concurrency to 1 (part 21 explains why).
KV fp8 and 4 streams: 293 tok/s combined, single stream still ~150
The weights got 9 GB smaller, which frees about 3 GB per card for the KV cache. I compared fp4 and fp8 KV at --max_batch 2, then raised concurrency with fp8. The single-stream column barely moves; watch the total.
| KV / streams | single code / math / prose | per stream | total | KV pool (tokens) |
|---|---|---|---|---|
| fp4 / 2 | 149.8 / 178.3 / 104.0 | 99.9 | 199.8 | 564,480 |
| fp8 / 2 | 151.0 / 178.3 / 101.5 | 108.3 | 216.6 | 321,152 |
| fp8 / 3 | 150.5 / 178.2 / 101.4 | 81.0 | 243.1 | 302,976 |
| fp8 / 4 | 150.6 / 178.3 / 101.5 | 73.4 | 293.4 | 257,920 |
fp8 KV is faster than fp4 once there's concurrency: two streams went from 99.9 to 108.3 each. My read is that fp4 KV has to be unpacked on every read, and that shows up when there's more KV to read.
The KV pool shrank from 564K to 258K tokens. The per-request limit, 257,920, is still above the FP8 build's 222,080.
After switching production over, I re-measured on the live service: single stream 151.7 / 178.8 / 101.9, two at once 107.8 each, four at once 74.3 each.
The cost: a 54K-token prompt prefills in 50.6 s instead of 40.5 s
Long prompts prefill 25% slower. Prefill processes thousands of tokens at once, so it's bound by compute rather than weight reads, and unpacking NVFP4 before the math is pure extra work when compute is the limit. Repeated prefixes still hit the prefix cache: sending the same 8K prefix a second time takes 1.8 s.
So this build suits mostly-short prompts with several concurrent users, like agents messaging each other. If you mostly feed it new documents of tens of thousands of tokens, the FP8 build is faster.
Capability: 131 / 128 on 140 questions, against 132 / 127 for FP8
Same 140 questions as before: GSM8K 50, MATH-500 30, MuSR 20, HumanEval+ 40. Compare the FP8 and self-made rows.
| GSM8K | MATH-500 | MuSR | HumanEval+ | total | |
|---|---|---|---|---|---|
| FP8, no thinking | 50 | 30 | 16 | 36 | 132 |
| self NVFP4, no thinking | 50 | 29 | 17 | 35 | 131 |
| FP8, thinking + low | 49 | 30 | 12 | 36 | 127 |
| self NVFP4, thinking + low | 50 | 30 | 12 | 36 | 128 |
On the same base model, going from FP8 to NVFP4 moved the score by one question in each mode, which is within noise. lyf's checkpoint scored 124 without thinking and 129 with thinking at low effort (full rows in the deep dive). Its −8 without thinking can't be split between quantization and the change of base model.
Deep dive: the shared global scale, bit-exact checks, the DFlash re-sweep and two 4-bit dead ends
You can skip this section without losing anything you need to run the setup. It holds the full data behind each claim above.
The shared global scale: 128 groups, and in_proj_a is the one that bites
The first version of the converter computed a global scale per tensor. The format and shapes were right, but the values didn't match lyf's. ModelOpt shares one global scale across a group of tensors, taken from the group's max absolute value:
- attention q/k/v: 16 layers → 16 groups
- GDN
in_proj_qkv/z/a/b: 48 layers → 48 groups - MLP gate/up: 64 layers → 64 groups
That's 128 groups. in_proj_a is the easiest to miss. Its own max is small, so computing its scale alone misses the larger values in the other three projections. Switching to per-group scales made the output bit-exact.
Bit-exact check against lyf's ModelOpt output
Method: decode lyf's NVFP4 weights to BF16, keeping FP4 negative zero, then re-quantize with the script and compare byte for byte.
| layer | packed weights | block scales | global scale |
|---|---|---|---|
Attention, layer 11 self_attn.q_proj | 31,457,280 / 31,457,280 bytes | 3,932,160 / 3,932,160 | 4 / 4 |
GDN, layer 0 linear_attn.out_proj | 15,728,640 / 15,728,640 | 1,966,080 / 1,966,080 | 4 / 4 |
MLP, layer 0 mlp.gate_proj | 44,564,480 / 44,564,480 | 5,570,560 / 5,570,560 | 4 / 4 |
For in_proj_qkv, re-estimating the global scale from the decoded BF16 can be off by 1 ULP, while the packed weights and block scales stay bit-identical. The decoded values can't fully recover the original BF16 extremes, so this is expected. I also re-quantized from the original BF16 for in_proj_a, q_proj and gate_proj in the final output: all three tensor kinds were bit-identical.
DFlash re-sweep on lyf's checkpoint: 8 draft tokens with NVFP4 draft quantization is still best
A faster main model could shift the best draft settings, so I re-swept on lyf's checkpoint.
| draft tokens / draft quant | code / math / prose | acceptance, positions 1–3 |
|---|---|---|
| 8 / nvfp4 | 136.0 / 185.6 / 102.5 | 80.8 / 59.1 / 44.7% |
| 8 / nvfp4_head | 118.6 / 174.4 / 94.8 | 84.5 / 63.2 / 46.8% |
| 8 / off | 123.8 / 162.6 / 92.5 | 83.6 / 65.6 / 50.4% |
| 6 / nvfp4 | 129.6 / 156.4 / 97.2 | 82.7 / 64.8 / 48.5% |
| 4 / nvfp4 | 113.2 / 124.7 / 97.3 | 85.4 / 68.1 / 51.2% |
Quantizing less of the draft raises acceptance but slows the draft down more than that gains. nvfp4_head (only the draft's output head quantized) is slower than no quantization at all; I didn't investigate why. Fewer draft tokens raise per-position acceptance but produce fewer tokens per step, and math takes the biggest hit. The checkpoint rejects values above 8.
KV and concurrency detail: acceptance, prefill and short-request latency per setting
| KV / streams | acceptance, positions 1–3 | 54K prompt prefill | short-request first token |
|---|---|---|---|
| fp4 / 2 | 81.9 / 63.1 / 47.8% | 49.9 s | 0.53 s |
| fp8 / 2 | 79.3 / 61.4 / 50.1% | 50.5 s | 0.47 s |
| fp8 / 3 | 80.3 / 62.8 / 51.7% | 50.6 s | 0.51 s |
| fp8 / 4 | 80.3 / 62.9 / 52.5% | 50.6 s | 0.53 s |
All four settings answered the 54K request correctly. The per-request limit stays at 262,144 for fp8/2 and fp8/3 and drops to 257,920 at fp8/4.
orcarouter's official NVFP4: loads fine, runs NVFP4 kernels, and is still slow
config.json has two quantization groups: 168 NVFP4 W4A4 targets (the MLP of the first 56 layers) and 232 FP8 W8A8 targets (attention q/k/v/o, GDN in_proj_qkv / in_proj_z / out_proj, and the last 8 layers' MLP); lm_head stays BF16. It loads without errors, the log shows NVFP4 prefill and decode kernels, it answers 17*23 = 391, and an image test comes back correct. It's just slow.
| checkpoint | code / math / prose | no-spec code | EN acceptance, positions 1–3 | KV pool (tokens) |
|---|---|---|---|---|
| FP8 | 116.6 / 135.8 / 74.6 | 32.9 | 82.7 / 67.3 / 52.1% | 222,080 |
| official NVFP4 | 65.3 / 77.9 / 43.6 | 32.4 | 86.3 / 68.3 / 55.5% | 447,104 |
| self NVFP4 | 149.9 / 178.0 / 104.3 | 45.1 | 88.8 / 73.4 / 55.0% | 564,480 |
Acceptance is higher than FP8 and no-spec speed matches FP8. It only falls behind when verifying several tokens at once.
Two 4-bit dead ends with FastLLM's own INT4
I also tried FastLLM's native INT4 format (int4g, groups of 128).
TP=2 has no INT4 matmul kernel. The model loads fine, 12 GB per card. On the first request, GPU utilization sits at 0% while CPU runs at 110%, and 16 tokens hadn't finished after 120 s. FastllmCudaMatMulFloatInt4Group exists only in the single-GPU cudadevice.cpp, not in multicudadevice.cpp, so tensor parallel silently falls back to CPU. The hung process ignores SIGTERM and needs kill -9.
Single-card INT4 is slow. Without speculation it runs 20.2 tok/s, which, at roughly 15 GB of weights on one card, works out to about 300 GB/s, half the 2080 Ti's bandwidth, so the unpacking is too expensive. Single card plus DFlash2 crashes with Resource deadlock avoided.
--dtype fp8_e4m3 doesn't convert a BF16 checkpoint
Passing --dtype fp8_e4m3 with a BF16 checkpoint doesn't produce FP8. FastLLM silently loads FP16 instead. For FP8 or NVFP4, convert offline first.
Pinned versions
- FastLLM: upstream
a2bf07fd+ revert ofd4b04876+ PR #749 + PR #756 (same as part 21) - Main model: orcarouter/Qwen3.8-27B-Uncensored BF16 → self-made NVFP4,
total_size20,558,935,392 bytes - Draft: z-lab/Qwen3.8-27B-DFlash2, quantized to NVFP4 at runtime (
FASTLLM_DRAFT_QUANT=nvfp4) - CUDA 12.4, sm_75, self-built NCCL 2.31.2
Limitations and next steps
- KV fp8 and 4 streams were set after the 140-question runs. I haven't re-run capability with them.
- The converter is only verified on Qwen3.8-27B's architecture.
- The 25% slower long-prompt prefill has no fix. It would take work on FastLLM's NVFP4 prefill kernel.
Also in this series:
FAQ
- Does NVFP4 speed up an RTX 2080 Ti even though the card has no FP4 hardware?
- Yes, in FastLLM. It keeps NVFP4 weights 4-bit in VRAM and unpacks them to FP16 inside the kernel, then runs FP16 tensor cores. Generation is bound by reading weights, and NVFP4 cuts the read from 29 GB to 20 GB per token. On two 2080 Tis, Qwen3.8-27B went from 32.9 to 45.1 tok/s on code without speculative decoding, and from 116.1 to 151.7 tok/s with the DFlash2 draft.
- Why is orcarouter's official Qwen3.8-27B NVFP4 slower than FP8 in FastLLM?
- It is mixed precision. Its config has 168 NVFP4 targets (the MLP of the first 56 layers) and 232 FP8 targets (attention and GDN projections, the last 8 layers' MLP); lm_head stays BF16. With DFlash2 it ran code at 65.3 tok/s against 116.6 for the FP8 checkpoint. A version I converted myself with every linear layer in NVFP4 ran at 149.9. Draft acceptance is higher than FP8, so my guess, not separately verified, is that the per-channel FP8 weights take a slow path when verifying several tokens at once.
- How do I convert a BF16 Qwen3.8-27B checkpoint to NVFP4 that FastLLM can read?
- Download the BF16 original (orcarouter/Qwen3.8-27B-Uncensored, about 56 GB), download only the json and recipe.yaml files from lyf/Qwen3.8-27B-Heretic-ARA-NVFP4-MTP-VL as a format reference, and run the numpy converter at ai-muninn.com/tools/qwen38-nvfp4-convert.py. No GPU and no calibration are needed. It quantizes every linear layer to NVFP4, shares one global scale across attention q/k/v, the four GDN input projections, and MLP gate/up, and keeps lm_head, embeddings, vision, conv1d and MTP in BF16. Output is 20 GB, and sampled layers are bit-identical to lyf's ModelOpt output.
- Does NVFP4 make Qwen3.8-27B dumber than FP8?
- Not measurably, when the base model is the same. On 140 questions (GSM8K, MATH-500, MuSR, HumanEval+), FP8 scored 132 without thinking and 127 with thinking at low effort; my NVFP4 conversion of the same base scored 131 and 128. lyf's NVFP4 scored 124 without thinking, but it uses a different uncensored base, so that drop can't be split between quantization and the base model.
- What does NVFP4 cost on a 2080 Ti?
- Long-prompt prefill. A 54K-token prompt took 50.6 s against 40.5 s with FP8, 25% slower. Prefill processes thousands of tokens at once and is limited by compute, so unpacking NVFP4 to FP16 is pure extra work there. Repeated prefixes hit the prefix cache: sending the same 8K prefix again took 1.8 s.
Read next
- 2026-09-28Dual RTX 2080 Ti FastLLM: 16% Faster Decoding and No More Head-of-Line Blocking
Qwen3.8-27B FP8 + DFlash2 on two modded 2080 Tis: code 100.3 → 116.1 tok/s, two requests at once, and a scheduler patch that stops long prompts blocking.
- 2026-09-21Two Modded RTX 2080 Tis Hit 153.8 tok/s on Qwen3.8-27B With FastLLM + DFlash2
FastLLM plus a DFlash2 draft runs Qwen3.8-27B FP8 at 153.8 tok/s on code across two modded 2080 Tis. Prose drops to 53.1. Full recipe and four traps.
- 2026-09-21Ternary Bonsai 2 on One RTX 2080 Ti: Qwen3.8-27B in 7.2 GB Scores 130 of 140
Ternary Bonsai 2 27B scored 130/140 on one modded 2080 Ti vs 131 for Qwen3.8-27B Q8_0 on two cards. Build steps, the flag that matters, why smaller is slower.
- 2026-08-28[Benchmark] A 177B MoE on Three Modded 2080 Tis: 23 tok/s at 128K, 77 on a File Edit
Qwen3.8-Flash-Next (176.94B, arch qwen4exp) runs on 3× modded 2080 Ti 22G at 23.14 tok/s at 128K context — and 3.51× that on a file edit, with no draft model.
Don't miss the next one
Subscribe, and you won't.
One-click unsubscribe anytime.