~/blog/fastllm-dual-2080ti-nvfp4-self-quant

改裝 2080 Ti 22G · part 22

NVFP4 Without FP4 Hardware: Qwen3.8-27B on Two RTX 2080 Tis at 151.7 tok/s

❯ cat --toc

TL;DR

I converted Qwen3.8-27B from BF16 to NVFP4 myself and serve it on two modded 2080 Ti 22G with FastLLM + DFlash2. Code went from 116.1 to 151.7 tok/s, math 135.5 → 178.8, prose 74.2 → 101.9, and the server now takes 4 requests at once (297 tok/s combined). The 2080 Ti has no FP4 hardware. It still wins because generation is bound by reading weights, and NVFP4 cuts that read from 29 GB to 20 GB per token. orcarouter's official NVFP4 is slower than FP8 here. Cost: a 54K-token prompt prefills 25% slower. 140-question score: 131/128 vs 132/127.

Modded 2080 Ti 22G series #22 cover: two dual-fan graphics cards side by side, a thick stack of data bricks pressed into thin slabs half the height, and the same glowing data stream passing through the cards faster

Same books, smaller editions: the model got faster by getting lighter

Movers bill by the trip, not by how many books you own. Put the same books in paperback and each box holds more pages, so the move takes fewer trips.

Text generation works the same way. Every token the model produces reads the whole set of weights out of VRAM once. With FP8 weights, Qwen3.8-27B reads 29 GB per token. With NVFP4 it reads 20 GB.

Part 21 left this machine running a self-built FastLLM with an FP8 model, two concurrent requests, and a scheduler that lets short requests cut in. The 2080 Ti is Turing and has no FP4 hardware, so I expected NVFP4 to be pointless on it. This part switches the main model to an NVFP4 checkpoint I converted myself from BF16, moves the KV cache to fp8, and raises concurrency to 4. Code went from 116.1 to 151.7 tok/s.

Self-made NVFP4 on two 2080 Tis: code 116.1 → 151.7 tok/s, 4 requests at once

Same two cards (TP=2), the same self-built FastLLM as part 21, the same benchmark script. Method: real chat requests, temperature 0, 512 output tokens, code / math / prose prompts, 3 runs each, median. Read the first row for speed and the last for whether anything broke.

FP8 (part 21)self-made NVFP4
code / math / prose tok/s116.1 / 135.5 / 74.2151.7 / 178.8 / 101.9
concurrent requests24 (297 tok/s combined)
main model size29 GB20 GB
54K-token prompt prefill40.5 s50.6 s
140 questions (no thinking / thinking+low)132 / 127131 / 128

With two requests at once, each stream went from about 88 to 107.8 tok/s. The per-request context limit went up from 222,080 to 257,920 tokens. When a short request arrives during a long prompt's prefill, its first token now takes 0.52 s instead of 1.18 s, and it finishes in 5.5 s (was 3.8 s). The 140 questions ran on the same NVFP4 model before the KV and concurrency changes.

Three changes produced this:

  1. Main model → NVFP4, converted myself from BF16, following a format that had already proven fast on this card.
  2. KV cache fp4 → fp8. The smaller weights freed enough VRAM, and fp8 KV is faster under concurrency.
  3. --max_batch 2 → 4.

Why NVFP4 helps without FP4 hardware: 29 GB → 20 GB read per token

Turing tensor cores do FP16, INT8 and INT4. They don't do FP4. FastLLM stores NVFP4 weights as they are, 4-bit in VRAM, and unpacks them to FP16 inside the kernel before the tensor cores touch them. The upstream README lists NVFP4 decode kernels tuned for SM75 (the 2080 Ti) for several matrices of width 5120, which is Qwen3.8-27B's hidden size.

So NVFP4 saves bytes read, not compute. That's the right thing to save here: generation is bottlenecked on reading weights, and the compute units sit mostly idle. Without speculative decoding, code went from 32.9 tok/s (FP8, 29 GB) to 45.1 tok/s (NVFP4, 20 GB), +37%.

Diagram: each generated token reads the full weights once. FP8 weights read 29 GB per token; NVFP4 weights read 20 GB, then pass through an unpack-to-FP16 step before the FP16 tensor cores. Without speculation, code runs 32.9 tok/s with FP8 and 45.1 tok/s with NVFP4.
Unpacking is extra work, but compute is idle during generation, so reading fewer bytes is what speeds it up: 32.9 → 45.1 tok/s without speculation.

DFlash2, the draft model from part 20, guesses several tokens ahead and lets the main model verify 8 positions per read of its weights. That multiplies whatever the base read speed is, so the gain carries over at a similar ratio: 116.1 → 151.7 on code.

Three NVFP4 checkpoints, more than 2× apart: 65.3, 136.0 and 149.9 tok/s on code

"NVFP4" on a model card doesn't tell you which layers are NVFP4. I tried three checkpoints under the same FastLLM and the same launch flags (KV fp4, --max_batch 2). Look at the code column and the last column together.

checkpointwhat's NVFP4code / math / proseno-spec code140 q
FP8 (part 21)none; linear layers FP8116.6 / 135.8 / 74.632.9132 / 127
orcarouter official NVFP4some MLP, rest FP865.3 / 77.9 / 43.632.4not tested
lyf Heretic-ARA NVFP4all linear layers136.0 / 185.5 / 102.145.3124 / 129
self-made orcarouter NVFP4all linear layers149.9 / 178.0 / 104.345.1131 / 128
Format map of three NVFP4 checkpoints. Official orcarouter: MLP layers 1 to 56 NVFP4; MLP 57 to 64 and the attention and GDN projections FP8; lm_head, embedding, vision, conv1d and MTP BF16. lyf and self-made: all MLP plus attention and GDN projections NVFP4; lm_head, embedding, vision, conv1d and MTP BF16. Code speed with DFlash2: 65.3, 136.0 and 149.9 tok/s.
Same label, different layouts. The official checkpoint keeps attention, GDN and the last 8 MLP layers in FP8 and runs code at 65.3 tok/s; the two all-linear NVFP4 checkpoints run at 136.0 and 149.9.

orcarouter's official NVFP4 comes from the same publisher as the uncensored FP8 I was already running, but it's mixed precision: the config lists 168 NVFP4 targets (MLP) and 232 FP8 targets. Without speculation it runs 32.4, about the same as FP8. With DFlash2 it drops to 65.3 on code, slower than the FP8 checkpoint. The draft isn't the problem: its acceptance is actually higher than with FP8. My guess, not separately verified: the per-channel FP8 weights take a slower path when the main model verifies 9 tokens at once.

lyf's checkpoint was converted with NVIDIA ModelOpt, every linear layer in NVFP4, and it's fast. But it's built on a different uncensored base (Heretic ARA), and its 140-question score without thinking drops to 124, mostly on MuSR and HumanEval+.

The self-made one takes orcarouter's BF16 original and converts it into lyf's format. It uses the same base model as my FP8 build, nearly matches lyf on speed, and scores about the same as FP8.

It's also a bit faster than lyf on code. Draft acceptance at positions 1–3 is 88.8 / 73.4 / 55.0%, against lyf's 82.8 / 64.3 / 52.4%. My guess: orcarouter's uncensoring is lighter, so the output stays closer to stock Qwen3.8, which is what the DFlash2 draft was trained on.

Convert BF16 to NVFP4 yourself: one numpy script, no GPU, no calibration

You need three things:

  • The BF16 original, orcarouter/Qwen3.8-27B-Uncensored, about 56 GB. It's gated: accept the terms on Hugging Face first.
  • The json and yaml files from lyf's repo, without the weights. The converter copies its tensor naming, shard layout and quantization_config.
  • The converter, /tools/qwen38-nvfp4-convert.py. It only needs numpy.
hf download orcarouter/Qwen3.8-27B-Uncensored --local-dir /mnt/nvme1t/models/orcarouter-qwen38-uncensored-bf16
hf download lyf/Qwen3.8-27B-Heretic-ARA-NVFP4-MTP-VL --include "*.json" "recipe.yaml" --local-dir ./nvfp4-ref
curl -O https://ai-muninn.com/tools/qwen38-nvfp4-convert.py
python3 qwen38-nvfp4-convert.py \
  /mnt/nvme1t/models/orcarouter-qwen38-uncensored-bf16 \
  /mnt/nvme1t/models/orcarouter-qwen38-uncensored-nvfp4-self \
  --reference ./nvfp4-ref

The output is 20 GB in three shards. What the script does:

  • Every linear layer → NVFP4. Values are packed as 4-bit E2M1, with one FP8 E4M3 scale per 16 values and one FP32 global scale per tensor.
  • Three groups share one global scale: attention q/k/v, the four GDN input projections (in_proj_qkv / z / a / b), and MLP gate/up. I only found this by comparing against lyf's weights. Computing each tensor's scale separately gives the right format with mismatched values.
  • Kept in BF16: lm_head, embed_tokens, the vision tower, conv1d, and MTP.
  • No calibration. ModelOpt also calibrates activation scales (input_global_scale) for W4A4. FastLLM on a 2080 Ti uses FP16 activations and never reads that field, so the script writes 1.0.

To check it, I decoded lyf's NVFP4 back to BF16 and re-quantized it with the script. For one attention layer, one GDN layer and one MLP layer, the packed weights, block scales and global scales were all bit-identical to lyf's. All 2,687 tensors in the output match lyf's in name, dtype, shape and shard.

Launch with KV fp8 and --max_batch 4, same FastLLM build

No rebuild. This is the production launch:

export LD_PRELOAD=/home/coolthor/nccl-install/lib/libnccl.so.2
export CUDA_VISIBLE_DEVICES=0,1
export FASTLLM_CUDA_GRAPH=0
export FASTLLM_DRAFT_QUANT=nvfp4
export FASTLLM_COOPERATIVE_LONG_PREFILL=1

~/venvs/ftllm-dual/bin/ftllm server \
  -p /mnt/nvme1t/models/orcarouter-qwen38-uncensored-nvfp4-self --model_name qwen38-27b-ud \
  --tp 2 --max_batch 4 --chunked_prefill_size 4096 \
  --gpu_mem_ratio 0.95 --kv_cache_dtype fp8_e4m3 \
  --speculative_algorithm dflash \
  --speculative_draft_model_path /mnt/nvme1t/models/dflash2-qwen38-zlab \
  --speculative_num_draft_tokens 8 \
  --prefix_cache true \
  --chat_template /mnt/nvme1t/models/qwen-fastllm-low/chat_template.jinja \
  --host 0.0.0.0 --port 8082

Leave out --tokens; with it, FastLLM clamps concurrency to 1 (part 21 explains why).

KV fp8 and 4 streams: 293 tok/s combined, single stream still ~150

The weights got 9 GB smaller, which frees about 3 GB per card for the KV cache. I compared fp4 and fp8 KV at --max_batch 2, then raised concurrency with fp8. The single-stream column barely moves; watch the total.

KV / streamssingle code / math / proseper streamtotalKV pool (tokens)
fp4 / 2149.8 / 178.3 / 104.099.9199.8564,480
fp8 / 2151.0 / 178.3 / 101.5108.3216.6321,152
fp8 / 3150.5 / 178.2 / 101.481.0243.1302,976
fp8 / 4150.6 / 178.3 / 101.573.4293.4257,920
Chart of per-stream and total decode speed for 1 to 4 concurrent streams with fp8 KV: one stream about 150 tok/s; two streams 108.3 each, 216.6 total; three 81.0 each, 243.1 total; four 73.4 each, 293.4 total.
Four streams add up to 293.4 tok/s, about 1.95× a single stream. Compute is mostly idle during generation, so each extra request barely adds weight reads.

fp8 KV is faster than fp4 once there's concurrency: two streams went from 99.9 to 108.3 each. My read is that fp4 KV has to be unpacked on every read, and that shows up when there's more KV to read.

The KV pool shrank from 564K to 258K tokens. The per-request limit, 257,920, is still above the FP8 build's 222,080.

After switching production over, I re-measured on the live service: single stream 151.7 / 178.8 / 101.9, two at once 107.8 each, four at once 74.3 each.

The cost: a 54K-token prompt prefills in 50.6 s instead of 40.5 s

Long prompts prefill 25% slower. Prefill processes thousands of tokens at once, so it's bound by compute rather than weight reads, and unpacking NVFP4 before the math is pure extra work when compute is the limit. Repeated prefixes still hit the prefix cache: sending the same 8K prefix a second time takes 1.8 s.

So this build suits mostly-short prompts with several concurrent users, like agents messaging each other. If you mostly feed it new documents of tens of thousands of tokens, the FP8 build is faster.

Capability: 131 / 128 on 140 questions, against 132 / 127 for FP8

Same 140 questions as before: GSM8K 50, MATH-500 30, MuSR 20, HumanEval+ 40. Compare the FP8 and self-made rows.

GSM8KMATH-500MuSRHumanEval+total
FP8, no thinking50301636132
self NVFP4, no thinking50291735131
FP8, thinking + low49301236127
self NVFP4, thinking + low50301236128

On the same base model, going from FP8 to NVFP4 moved the score by one question in each mode, which is within noise. lyf's checkpoint scored 124 without thinking and 129 with thinking at low effort (full rows in the deep dive). Its −8 without thinking can't be split between quantization and the change of base model.


Deep dive: the shared global scale, bit-exact checks, the DFlash re-sweep and two 4-bit dead ends

You can skip this section without losing anything you need to run the setup. It holds the full data behind each claim above.

The shared global scale: 128 groups, and in_proj_a is the one that bites

The first version of the converter computed a global scale per tensor. The format and shapes were right, but the values didn't match lyf's. ModelOpt shares one global scale across a group of tensors, taken from the group's max absolute value:

  • attention q/k/v: 16 layers → 16 groups
  • GDN in_proj_qkv / z / a / b: 48 layers → 48 groups
  • MLP gate/up: 64 layers → 64 groups

That's 128 groups. in_proj_a is the easiest to miss. Its own max is small, so computing its scale alone misses the larger values in the other three projections. Switching to per-group scales made the output bit-exact.

Bit-exact check against lyf's ModelOpt output

Method: decode lyf's NVFP4 weights to BF16, keeping FP4 negative zero, then re-quantize with the script and compare byte for byte.

layerpacked weightsblock scalesglobal scale
Attention, layer 11 self_attn.q_proj31,457,280 / 31,457,280 bytes3,932,160 / 3,932,1604 / 4
GDN, layer 0 linear_attn.out_proj15,728,640 / 15,728,6401,966,080 / 1,966,0804 / 4
MLP, layer 0 mlp.gate_proj44,564,480 / 44,564,4805,570,560 / 5,570,5604 / 4

For in_proj_qkv, re-estimating the global scale from the decoded BF16 can be off by 1 ULP, while the packed weights and block scales stay bit-identical. The decoded values can't fully recover the original BF16 extremes, so this is expected. I also re-quantized from the original BF16 for in_proj_a, q_proj and gate_proj in the final output: all three tensor kinds were bit-identical.

DFlash re-sweep on lyf's checkpoint: 8 draft tokens with NVFP4 draft quantization is still best

A faster main model could shift the best draft settings, so I re-swept on lyf's checkpoint.

draft tokens / draft quantcode / math / proseacceptance, positions 1–3
8 / nvfp4136.0 / 185.6 / 102.580.8 / 59.1 / 44.7%
8 / nvfp4_head118.6 / 174.4 / 94.884.5 / 63.2 / 46.8%
8 / off123.8 / 162.6 / 92.583.6 / 65.6 / 50.4%
6 / nvfp4129.6 / 156.4 / 97.282.7 / 64.8 / 48.5%
4 / nvfp4113.2 / 124.7 / 97.385.4 / 68.1 / 51.2%

Quantizing less of the draft raises acceptance but slows the draft down more than that gains. nvfp4_head (only the draft's output head quantized) is slower than no quantization at all; I didn't investigate why. Fewer draft tokens raise per-position acceptance but produce fewer tokens per step, and math takes the biggest hit. The checkpoint rejects values above 8.

KV and concurrency detail: acceptance, prefill and short-request latency per setting

KV / streamsacceptance, positions 1–354K prompt prefillshort-request first token
fp4 / 281.9 / 63.1 / 47.8%49.9 s0.53 s
fp8 / 279.3 / 61.4 / 50.1%50.5 s0.47 s
fp8 / 380.3 / 62.8 / 51.7%50.6 s0.51 s
fp8 / 480.3 / 62.9 / 52.5%50.6 s0.53 s

All four settings answered the 54K request correctly. The per-request limit stays at 262,144 for fp8/2 and fp8/3 and drops to 257,920 at fp8/4.

orcarouter's official NVFP4: loads fine, runs NVFP4 kernels, and is still slow

config.json has two quantization groups: 168 NVFP4 W4A4 targets (the MLP of the first 56 layers) and 232 FP8 W8A8 targets (attention q/k/v/o, GDN in_proj_qkv / in_proj_z / out_proj, and the last 8 layers' MLP); lm_head stays BF16. It loads without errors, the log shows NVFP4 prefill and decode kernels, it answers 17*23 = 391, and an image test comes back correct. It's just slow.

checkpointcode / math / proseno-spec codeEN acceptance, positions 1–3KV pool (tokens)
FP8116.6 / 135.8 / 74.632.982.7 / 67.3 / 52.1%222,080
official NVFP465.3 / 77.9 / 43.632.486.3 / 68.3 / 55.5%447,104
self NVFP4149.9 / 178.0 / 104.345.188.8 / 73.4 / 55.0%564,480

Acceptance is higher than FP8 and no-spec speed matches FP8. It only falls behind when verifying several tokens at once.

Two 4-bit dead ends with FastLLM's own INT4

I also tried FastLLM's native INT4 format (int4g, groups of 128).

TP=2 has no INT4 matmul kernel. The model loads fine, 12 GB per card. On the first request, GPU utilization sits at 0% while CPU runs at 110%, and 16 tokens hadn't finished after 120 s. FastllmCudaMatMulFloatInt4Group exists only in the single-GPU cudadevice.cpp, not in multicudadevice.cpp, so tensor parallel silently falls back to CPU. The hung process ignores SIGTERM and needs kill -9.

Single-card INT4 is slow. Without speculation it runs 20.2 tok/s, which, at roughly 15 GB of weights on one card, works out to about 300 GB/s, half the 2080 Ti's bandwidth, so the unpacking is too expensive. Single card plus DFlash2 crashes with Resource deadlock avoided.

--dtype fp8_e4m3 doesn't convert a BF16 checkpoint

Passing --dtype fp8_e4m3 with a BF16 checkpoint doesn't produce FP8. FastLLM silently loads FP16 instead. For FP8 or NVFP4, convert offline first.

Pinned versions

  • FastLLM: upstream a2bf07fd + revert of d4b04876 + PR #749 + PR #756 (same as part 21)
  • Main model: orcarouter/Qwen3.8-27B-Uncensored BF16 → self-made NVFP4, total_size 20,558,935,392 bytes
  • Draft: z-lab/Qwen3.8-27B-DFlash2, quantized to NVFP4 at runtime (FASTLLM_DRAFT_QUANT=nvfp4)
  • CUDA 12.4, sm_75, self-built NCCL 2.31.2

Limitations and next steps

  • KV fp8 and 4 streams were set after the 140-question runs. I haven't re-run capability with them.
  • The converter is only verified on Qwen3.8-27B's architecture.
  • The 25% slower long-prompt prefill has no fix. It would take work on FastLLM's NVFP4 prefill kernel.

Also in this series:

FAQ

Does NVFP4 speed up an RTX 2080 Ti even though the card has no FP4 hardware?
Yes, in FastLLM. It keeps NVFP4 weights 4-bit in VRAM and unpacks them to FP16 inside the kernel, then runs FP16 tensor cores. Generation is bound by reading weights, and NVFP4 cuts the read from 29 GB to 20 GB per token. On two 2080 Tis, Qwen3.8-27B went from 32.9 to 45.1 tok/s on code without speculative decoding, and from 116.1 to 151.7 tok/s with the DFlash2 draft.
Why is orcarouter's official Qwen3.8-27B NVFP4 slower than FP8 in FastLLM?
It is mixed precision. Its config has 168 NVFP4 targets (the MLP of the first 56 layers) and 232 FP8 targets (attention and GDN projections, the last 8 layers' MLP); lm_head stays BF16. With DFlash2 it ran code at 65.3 tok/s against 116.6 for the FP8 checkpoint. A version I converted myself with every linear layer in NVFP4 ran at 149.9. Draft acceptance is higher than FP8, so my guess, not separately verified, is that the per-channel FP8 weights take a slow path when verifying several tokens at once.
How do I convert a BF16 Qwen3.8-27B checkpoint to NVFP4 that FastLLM can read?
Download the BF16 original (orcarouter/Qwen3.8-27B-Uncensored, about 56 GB), download only the json and recipe.yaml files from lyf/Qwen3.8-27B-Heretic-ARA-NVFP4-MTP-VL as a format reference, and run the numpy converter at ai-muninn.com/tools/qwen38-nvfp4-convert.py. No GPU and no calibration are needed. It quantizes every linear layer to NVFP4, shares one global scale across attention q/k/v, the four GDN input projections, and MLP gate/up, and keeps lm_head, embeddings, vision, conv1d and MTP in BF16. Output is 20 GB, and sampled layers are bit-identical to lyf's ModelOpt output.
Does NVFP4 make Qwen3.8-27B dumber than FP8?
Not measurably, when the base model is the same. On 140 questions (GSM8K, MATH-500, MuSR, HumanEval+), FP8 scored 132 without thinking and 127 with thinking at low effort; my NVFP4 conversion of the same base scored 131 and 128. lyf's NVFP4 scored 124 without thinking, but it uses a different uncensored base, so that drop can't be split between quantization and the base model.
What does NVFP4 cost on a 2080 Ti?
Long-prompt prefill. A 54K-token prompt took 50.6 s against 40.5 s with FP8, 25% slower. Prefill processes thousands of tokens at once and is limited by compute, so unpacking NVFP4 to FP16 is pure extra work there. Repeated prefixes hit the prefix cache: sending the same 8K prefix again took 1.8 s.

Read next

Don't miss the next one

Subscribe, and you won't.

One-click unsubscribe anytime.