[Announcement] On hiatus until August 21

The open models are landing thick and fast right now, and my gallbladder picked exactly this moment to give out. Nothing new here for a couple of weeks.
Hi everyone. This is a genuinely great stretch to be into local models — the full DeepSeek V4 Flash, Qwen 3.8, MiniMax H3, all within days of each other. There is a ridiculous amount to play with. Go enjoy it.
I'd love to be joining in. Instead I've spent about a week in a hospital bed: gallstones, a blocked duct, and the kind of inflammation that gets you admitted rather than sent home with painkillers. I kept telling myself I could still get some benchmarking done between the bad stretches.

I can't even.
(Anyone who can actually run benchmarks through this is built differently. XD)
Look after yourselves out there.
Read next
- 2026-09-28[DGX Spark] Qwen3.8-Flash-Next on two DGX Sparks: 51.9 to 72.0 tok/s, and a draft vocabulary built from my own text
Sixteen days after Part 46, the same 40 prompts go from 51.9 to 72.0 tok/s with no score loss. The biggest gain: an MTP draft vocabulary built from my own Chinese text. Multi-turn TTFT fell from 3.30 s to 0.23 s.
- 2026-09-28NVFP4 Without FP4 Hardware: Qwen3.8-27B on Two RTX 2080 Tis at 151.7 tok/s
No FP4 hardware, still faster: a self-converted NVFP4 Qwen3.8-27B on two 2080 Tis runs code at 151.7 tok/s (was 116.1) and serves 4 requests at once.
- 2026-09-28Dual RTX 2080 Ti FastLLM: 16% Faster Decoding and No More Head-of-Line Blocking
Qwen3.8-27B FP8 + DFlash2 on two modded 2080 Tis: code 100.3 → 116.1 tok/s, two requests at once, and a scheduler patch that stops long prompts blocking.
- 2026-09-21Two Modded RTX 2080 Tis Hit 153.8 tok/s on Qwen3.8-27B With FastLLM + DFlash2
FastLLM plus a DFlash2 draft runs Qwen3.8-27B FP8 at 153.8 tok/s on code across two modded 2080 Tis. Prose drops to 53.1. Full recipe and four traps.
Don't miss the next one
Subscribe, and you won't.
One-click unsubscribe anytime.