[Announcement] On hiatus until August 21

The open models are landing thick and fast right now, and my gallbladder picked exactly this moment to give out. Nothing new here for a couple of weeks.
Hi everyone. This is a genuinely great stretch to be into local models — the full DeepSeek V4 Flash, Qwen 3.8, MiniMax H3, all within days of each other. There is a ridiculous amount to play with. Go enjoy it.
I'd love to be joining in. Instead I've spent about a week in a hospital bed: gallstones, a blocked duct, and the kind of inflammation that gets you admitted rather than sent home with painkillers. I kept telling myself I could still get some benchmarking done between the bad stretches.

I can't even.
(Anyone who can actually run benchmarks through this is built differently. XD)
Look after yourselves out there.
Read next
- 2026-09-11[Dev Workflow] Zero-Shot Voice Cloning on a MacBook: 24 Seconds In, Faster Than Realtime Out
Preset TTS voices don't sound like you. Zero-shot cloning with Qwen3-TTS and mlx-audio needs 24 seconds of reference audio and ten lines of Python, all local.
- 2026-09-11[Benchmark] LTX-2.5 vs MiniMax-H3 on one RTX 5090: 29s vs 81s for the same Chinese dialogue clip
LTX-2.5 and MiniMax-H3 on one RTX 5090, same Simplified Chinese line: 28.67s vs 81.22s end to end, both at CER 0%. Files, params, VRAM ceiling.
- 2026-09-06[Benchmark] Qwen3.8-Flash-Next NVFP4 on a DGX Spark: 41.7 tok/s, RAM for Traffic, Disk for the Dictionary
NVIDIA's NVFP4 checkpoint at 41.7 tok/s on one DGX Spark via nine bind-mounted vLLM files, plus six figures on why the 47.68 GiB n-gram table lives on NVMe.
- 2026-09-06[Benchmark] Two characters in one shot: MiniMax-H3 Ref2VA takes multiple reference images
Ref2VA is the only MiniMax-H3 mode that takes several reference images. Two characters, 243 frames, 437 s on one RTX 5090, and CER 6.8% on the dialogue.
Don't miss the next one
Subscribe, and you won't.
One-click unsubscribe anytime.