改裝 2080 Ti 22G · part 23
Qwen3.8-Flash-Next 177B on One RTX 2080 Ti: Strata Runs Code at 70 tok/s
❯ cat --toc
- Why I tried this: the 27B is maxed out, and the 177B used to need the whole machine
- One card runs code at 70.3 tok/s, 23% slower than two
- How Strata fits 177B on one card: the GPU holds what every token uses, RAM holds the experts
- What you need: a 22 GB card, 64 GB of RAM at minimum, about 90 GB of disk
- Install: one setup.py run builds the CUDA kernels and fetches the weights
- Config, launch, first request: an OpenAI-compatible server on port 8096
- Running next to the 27B costs Strata 1.3% and the 27B 7.7%
- Sharing the card with ComfyUI: 5 minutes to switch to Shio, 12 seconds back
- Deep dive: parallel, thinking effort, 16K numbers, dead ends and versions
- `parallel: 2` slows even a single request
- Thinking defaults to xhigh; low scored higher in under a third of the time
- The 16K numbers were measured at 32K max context, one shot each
- Dual-card speed tricks that didn't help
- Versions
- What I didn't measure
TL;DR
One modded RTX 2080 Ti 22GB plus 128 GB of RAM runs the 177B Qwen3.8-Flash-Next (IQ3_S) under Strata at 70.3 tok/s on code, 52.9 on Chinese, 51.5 on English. Two cards do 91.4 on code; one card is 23% slower. On the GPU, Strata uses 21.6 GB of the card's 22 GB of VRAM for attention, the router, shared experts, the KV cache, and the most-used experts. System RAM holds a copy of every expert, and the CPU runs the experts that aren't on the GPU directly from system RAM; available RAM drops by 37–48 GB once Strata is up. It runs next to my 27B with a 1.3% slowdown and swaps with ComfyUI in about 5 minutes. Caveat: Strata defaults to xhigh thinking effort; send low.

Why I tried this: the 27B is maxed out, and the 177B used to need the whole machine
My ai-lab box is an EPYC with 128 GB of RAM and three modded RTX 2080 Ti 22GB. Cards 0 and 1 run Qwen3.8-27B under FastLLM (part 22). Card 2 belongs to ComfyUI for image and video work, which comes in bursts and sits idle for days at a time.
The 27B is about as fast as I can make it, but it's still a 27B. I wanted the 177B Qwen3.8-Flash-Next. Part 17 ran the same 177B with llama.cpp across all three cards: 23.14 tok/s at 128K context (26.16 at 32K). That took the whole machine. The 27B and ComfyUI had to stop.
This part puts the 177B on the idle third card alone, using Strata, and leaves the 27B running.
One card runs code at 70.3 tok/s, 23% slower than two
Same prompts on one card (GPU2) and on two (GPU0+1). Method: streaming, thinking off, 650 max tokens, 3 runs each, median; decode speed comes from Strata's /metrics.
| one card | two cards | change | |
|---|---|---|---|
| code / Chinese / English tok/s | 70.3 / 52.9 / 51.5 | 91.4 / 60.0 / 60.0 | −23% / −12% / −14% |
| 16K prompt: time to first token | 12.2 s | 10.6 s | |
| 16K prompt: prefill | 1,325 tok/s | 1,525 tok/s | |
| 16K prompt: decode | 49.8 tok/s | 66.1 tok/s |
VRAM in use on the single card: 21.6 GB.
Code generation takes the biggest hit without the second card. My guess is that code hits a concentrated set of experts, so the second card's extra cache catches more of them. I didn't measure expert usage to confirm that.
Part 17's 77.56 tok/s doesn't belong in this comparison. That was n-gram speculative decoding on a file-edit task where the output is mostly copied from the input.
How Strata fits 177B on one card: the GPU holds what every token uses, RAM holds the experts
Flash-Next is a mixture-of-experts model: 48 layers, 512 experts each. Each token runs through 10 routed experts plus 1 shared expert per layer, so most of the weights sit unused on any given token. Strata splits the model along that line.

- On the GPU: the attention and DeltaNet mixers, the router, the shared experts, the output head, the MTP draft layer and the KV cache (with 128K context, only the most-read 32K stays here). Whatever VRAM is left becomes a cache of the most-used experts. The cache adapts as you chat.
- In RAM: all 24,576 routed experts (48 × 512), pinned. When a token needs an expert that isn't cached, the CPU computes it in place, in parallel with the GPU, instead of copying it over.
- The n-gram table: 28.8 GB. By default Strata reads it directly from SSD (
--ple-io direct). I set--ple-io ram; I measured no speed difference. - MTP: the model ships with its own multi-token prediction layer. It drafts a few tokens ahead and the main model checks them in one pass, so the output is identical to running without it.
What you need: a 22 GB card, 64 GB of RAM at minimum, about 90 GB of disk
- GPU: I used a modded 2080 Ti 22GB, and Strata used 21.6 GB of it. Strata compiled cleanly for sm_75 (Turing).
- RAM: Strata's setup lists IQ3_S at 62 GB and says it "needs a 64 GB PC with little else running". On my box, available memory dropped by 47.6, 37.0 and 39.5 GiB across three loads, with the n-gram table in RAM. That's the whole machine's available memory, not Strata's own footprint.
- Disk: the two GGUF shards are 54.8 GB + 28.8 GB = 83.6 GB. Add the MTP weights setup fetches (6.5 GiB) and the pack (1.5 GiB), and
duputs the whole directory at about 86 GiB. Leave 90 GB. - Toolchain: Python 3.10+, g++ and nvcc. No sudo needed if those are already there. I built against CUDA 12.4.
The model is ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF, quant IQ3_S.
Install: one setup.py run builds the CUDA kernels and fetches the weights
Clone the repo and pin the commit I used:
git clone https://github.com/Niko1221/Strata
cd Strata
git checkout 6f32ec07
Then create a venv in the repo and run setup. I kept every cache on the NVMe drive:
python3 -m venv .venv
export XDG_CONFIG_HOME=/mnt/nvme1t/strata/config
export PIP_CACHE_DIR=/mnt/nvme1t/strata/pip-cache
export TMPDIR=/mnt/nvme1t/strata/tmp
export HF_HOME=/mnt/nvme1t/strata/data/hf
.venv/bin/python setup.py --family qwen --model IQ3_S --gpu 0 --cuda 12 --build \
--context 32768 --kv int8 --vision no \
--data-dir /mnt/nvme1t/strata/data \
--port 8096 --host 127.0.0.1 --no-browser --no-start \
--low-ram off --yes
A few notes:
--build --cuda 12compiles from source.cuobjdumpon the result shows 58 sm_75 ELF sections, so the 2080 Ti gets native kernels.--low-ram offwas needed because the 27B and ComfyUI were running during install, so setup saw too little free RAM.- Setup fetches the MTP weights from Qwen's official repo and checks each tensor's SHA-256.
- I pre-downloaded the two shards with
curl -C -(resumable) and checked their SHA-256 against Hugging Face. - It's done when setup prints
All set.
Config, launch, first request: an OpenAI-compatible server on port 8096
My config file, strata-iq3_s-gpu2.json, sets these flags:
| flag | value | what it does |
|---|---|---|
--expert-cache / --prefill | auto | expert cache size and prefill mode chosen by Strata |
--spec / --spec-min-p | 4 / 0.5 | speculative drafting settings for the MTP layer |
--max-context / --kv | 131072 / int8 | 128K context, int8 KV cache |
--kv-resident | 32768 | keeps the most-read 32K of KV on the GPU; the rest streams from RAM |
--ple-io | ram | n-gram table in RAM instead of SSD reads |
It also sets --mtp to the MTP weights directory, "parallel": 1 and the alias shio. Keep parallel at 1; the deep dive explains why.
To pin Strata to the third card, select it by UUID. Strata still sees it as GPU 0:
export CUDA_VISIBLE_DEVICES=<GPU UUID>
export STRATA_BF16_TC=1
.venv/bin/python serve/server.py --engine strata \
--config strata-iq3_s-gpu2.json \
--host 0.0.0.0 --port 8096 --gpu 0
Loading takes a few minutes. It's ready when /v1/models returns 200. Then:
curl -s http://localhost:8096/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "shio",
"messages": [{"role": "user", "content": "Write a Python function that merges two sorted lists."}],
"reasoning_effort": "low",
"stream": false
}'
Keep "reasoning_effort": "low". Strata turns thinking on at effort xhigh by default, which on a 140-question set took 3.4× as long for a lower score. Details are in the deep dive.
Running next to the 27B costs Strata 1.3% and the 27B 7.7%
The two models are on different GPUs, so the only shared resources are the CPU and memory bandwidth. I sent the same code prompt to both at once, 3 rounds.
| alone | both at once | change | |
|---|---|---|---|
| Strata (one card) | 70.3 tok/s | 69.4 | −1.3% |
| 27B (two cards) | 243.1 tok/s | 224.4 | −7.7% |
I'm assuming the contention is CPU or memory bandwidth; I didn't isolate it. Also, the overlap is short: the 27B finishes in about 3 s and Strata takes 10–12 s, so they only compete during Strata's first ~3 s.
The 27B never restarted. A health check every 5 s returned 106/106 HTTP 200.
The 243.1 uses this article's method (streaming, thinking off, 650 tokens). Part 22's 151.7 used a different method; don't compare them.
Sharing the card with ComfyUI: 5 minutes to switch to Shio, 12 seconds back
I call the Strata service on the third GPU Shio, after 汐 (tide). GPU2 runs Shio normally and switches to ComfyUI when I do video work.
I set Conflicts= between the two systemd user services, so starting one stops the other. A small script drives them:
~/bin/gpu2-mode shio # switch to Strata
~/bin/gpu2-mode comfy # switch to ComfyUI
~/bin/gpu2-mode status
The script has three guards:
- It refuses to switch to Shio while ComfyUI's queue has jobs.
- It waits for Shio to go idle before switching to ComfyUI.
- If the target side isn't ready within 10 minutes, it rolls back to the previous side.
Measured over 3 rounds:
| direction | ready | first result |
|---|---|---|
| ComfyUI → Shio | 281–334 s | first request 2.6–3.4 s after ready |
| Shio → ComfyUI | 12.4 s | first 512px image 30.5–31 s (cold model load); second 18.3–18.4 s |
From the command to the first ComfyUI image is about 43 s. Rounds 2 and 3 were faster than round 1, possibly from the OS page cache; I didn't isolate that. ComfyUI is still the boot default; Shio isn't enabled at boot.
Deep dive: parallel, thinking effort, 16K numbers, dead ends and versions
You can skip this section; everything you need to run the setup is above. This is the data behind the settings.
parallel: 2 slows even a single request
Strata's batch path doesn't use MTP drafting, but its docs say a request running alone still takes the MTP path. I measured a slowdown anyway, even with one request, and haven't tracked down why. Setting parallel back to 1 fixes it.
| setting | requests | per request | total |
|---|---|---|---|
parallel: 2 | 1 | 60.32 tok/s | 54.49 (over full batch wall time) |
parallel: 2 | 2 | 30.39 / 32.26 | 60.23 |
parallel: 1 | 1 | 67.2, 70.1 |
Two concurrent requests together get less throughput than one request with parallel: 1. The ComfyUI switching test happened to run with parallel: 2 and saw 44.4, 60.3 and 44.8 tok/s. Back on parallel: 1, two checks gave 67.2 and 70.1.
Thinking defaults to xhigh; low scored higher in under a third of the time
Strata enables thinking at effort xhigh by default. I don't have a capability run on Shio itself, but I have the same 140-question set on the same model, Qwen3.8-Flash-Next, on a different machine and engine (GX10, vLLM, NVFP4):
| effort | score | time |
|---|---|---|
| low | 126 | 1,736 s |
| xhigh | 117 | 5,902 s |
Only the relative gap carries over.
Two ways to fix it:
- Send
"reasoning_effort": "low"on every request. - What I did: 3 lines in
serve/server.py's OpenAI entry point. If a request carries no effort at all, fill inlow.
Strata's /settings page has a shared default, but it only applies when the request has no thinking parameters at all. Clients that send only an "enable thinking" flag skip it. My client, Pi, does exactly that, so the server patch was the only thing that worked for it.
The 16K numbers were measured at 32K max context, one shot each
The 16K-prompt rows (TTFT, prefill, decode) ran with --max-context 32768, one shot per setting. Production uses 128K with --kv-resident 32768. I re-measured code on the production config: 70.1 tok/s, in line with the 70.3 from the main table.
Dual-card speed tricks that didn't help
I tried several things on the two-card setup; the next post covers them in detail.
--ple-io ramvsdirect: no difference.--peer-device(use NVLink between the cards): about 10% slower.- Using the third card as an expert helper for the two-card setup: code went from 91.4 to 48 tok/s (−47%).
The one flag that did help was STRATA_BF16_TC=1. On two cards it cut 16K TTFT from 12.6 s to 10.6 s. Single-card Shio uses that and nothing else from this list.
Versions
- Strata engine 0.1.39, commit
6f32ec07, built with--build --cuda 12 - CUDA 12.4, sm_75;
cuobjdumpshows 58 sm_75 ELF sections - Model: ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF, IQ3_S, two shards (54.8 GB + 28.8 GB)
- MTP weights from Qwen's official repo, per-tensor SHA-256 checked
- Draft vocab: Strata's built-in CJK subset, 106,299 ids
- Env:
STRATA_BF16_TC=1
What I didn't measure
- The 140-question set on Shio itself.
Also in this series:
FAQ
- Can a single RTX 2080 Ti run the 177B Qwen3.8-Flash-Next?
- Yes, with Strata and enough system RAM. One modded 2080 Ti 22GB with 128 GB of RAM runs the IQ3_S quant at 70.3 tok/s on code, 52.9 on Chinese and 51.5 on English (streaming, thinking off, median of 3). The card uses 21.6 GB of VRAM. Two cards run the same prompts at 91.4, 60.0 and 60.0.
- How does Strata fit a 177B MoE model on a 22 GB GPU?
- The GPU holds the parts every token uses: attention and DeltaNet mixers, the router, shared experts, the output head, the MTP draft layer and the KV cache (at 128K context, only the most-read 32K of it). The rest of VRAM becomes a cache of the most-used experts, which adapts as you chat. All 24,576 routed experts sit pinned in system RAM, and the CPU computes any expert that is not cached, in parallel with the GPU.
- How much system RAM does Strata need for Qwen3.8-Flash-Next IQ3_S?
- Strata's setup lists IQ3_S at 62 GB and says it needs a 64 GB PC with little else running. On my 128 GB box, available system memory dropped by 37-48 GB across three loads, with the 28.8 GB n-gram table loaded into RAM.
- Why is Strata slow or verbose on simple prompts?
- Strata turns thinking on at effort xhigh by default. Send reasoning_effort low with each request, or patch the server to fill in low when a request carries no effort. On a 140-question set with the same model on other hardware, low scored 126 in 1,736 s and xhigh scored 117 in 5,902 s.
- Can Strata and another model share the same machine?
- On different GPUs, yes. With Strata on one 2080 Ti and a 27B on the other two, sending the same code prompt to both at once slowed Strata from 70.3 to 69.4 tok/s and the 27B from 243.1 to 224.4. The 27B finishes in about 3 s, so the overlap only covers the start of Strata's answer.
Read next
- 2026-09-21Ternary Bonsai 2 on One RTX 2080 Ti: Qwen3.8-27B in 7.2 GB Scores 130 of 140
Ternary Bonsai 2 27B scored 130/140 on one modded 2080 Ti vs 131 for Qwen3.8-27B Q8_0 on two cards. Build steps, the flag that matters, why smaller is slower.
- 2026-08-22[Benchmark] Two Modded 2080 Tis Reach 59.6 tok/s on Qwen3.8-27B With llama.cpp Tensor Parallel
Two modded 22GB RTX 2080 Tis hit 59.6 tok/s on Qwen3.8-27B using llama.cpp's -sm tensor split mode plus MTP, after -sm row was deleted upstream. Config, gotchas and the failed routes included.
- 2026-08-21[Benchmark] Free 15% Speedup on a 2080 Ti: One Broken Chat Template, One Starving MTP Head
A frozen GGUF chat template was killing my KV cache; a community fix plus MTP draft depth 4 took Qwen3.8-27B from 41.6 to 47.6 tok/s on a 2080 Ti 22GB.
- 2026-09-28NVFP4 Without FP4 Hardware: Qwen3.8-27B on Two RTX 2080 Tis at 151.7 tok/s
No FP4 hardware, still faster: a self-converted NVFP4 Qwen3.8-27B on two 2080 Tis runs code at 151.7 tok/s (was 116.1) and serves 4 requests at once.
Don't miss the next one
Subscribe, and you won't.
One-click unsubscribe anytime.