~/blog/strata-flashnext-one-2080ti

改裝 2080 Ti 22G · part 23

Qwen3.8-Flash-Next 177B on One RTX 2080 Ti: Strata Runs Code at 70 tok/s

❯ cat --toc

TL;DR

One modded RTX 2080 Ti 22GB plus 128 GB of RAM runs the 177B Qwen3.8-Flash-Next (IQ3_S) under Strata at 70.3 tok/s on code, 52.9 on Chinese, 51.5 on English. Two cards do 91.4 on code; one card is 23% slower. On the GPU, Strata uses 21.6 GB of the card's 22 GB of VRAM for attention, the router, shared experts, the KV cache, and the most-used experts. System RAM holds a copy of every expert, and the CPU runs the experts that aren't on the GPU directly from system RAM; available RAM drops by 37–48 GB once Strata is up. It runs next to my 27B with a 1.3% slowdown and swaps with ComfyUI in about 5 minutes. Caveat: Strata defaults to xhigh thinking effort; send low.

Modded 2080 Ti 22G series #23 cover: a single dual-fan graphics card carrying a giant mixture-of-experts model, with most experts resting in a large bank of system memory beside it

Why I tried this: the 27B is maxed out, and the 177B used to need the whole machine

My ai-lab box is an EPYC with 128 GB of RAM and three modded RTX 2080 Ti 22GB. Cards 0 and 1 run Qwen3.8-27B under FastLLM (part 22). Card 2 belongs to ComfyUI for image and video work, which comes in bursts and sits idle for days at a time.

The 27B is about as fast as I can make it, but it's still a 27B. I wanted the 177B Qwen3.8-Flash-Next. Part 17 ran the same 177B with llama.cpp across all three cards: 23.14 tok/s at 128K context (26.16 at 32K). That took the whole machine. The 27B and ComfyUI had to stop.

This part puts the 177B on the idle third card alone, using Strata, and leaves the 27B running.

One card runs code at 70.3 tok/s, 23% slower than two

Same prompts on one card (GPU2) and on two (GPU0+1). Method: streaming, thinking off, 650 max tokens, 3 runs each, median; decode speed comes from Strata's /metrics.

one cardtwo cardschange
code / Chinese / English tok/s70.3 / 52.9 / 51.591.4 / 60.0 / 60.0−23% / −12% / −14%
16K prompt: time to first token12.2 s10.6 s
16K prompt: prefill1,325 tok/s1,525 tok/s
16K prompt: decode49.8 tok/s66.1 tok/s

VRAM in use on the single card: 21.6 GB.

Code generation takes the biggest hit without the second card. My guess is that code hits a concentrated set of experts, so the second card's extra cache catches more of them. I didn't measure expert usage to confirm that.

Part 17's 77.56 tok/s doesn't belong in this comparison. That was n-gram speculative decoding on a file-edit task where the output is mostly copied from the input.

How Strata fits 177B on one card: the GPU holds what every token uses, RAM holds the experts

Flash-Next is a mixture-of-experts model: 48 layers, 512 experts each. Each token runs through 10 routed experts plus 1 shared expert per layer, so most of the weights sit unused on any given token. Strata splits the model along that line.

Diagram of Strata's placement on one 2080 Ti 22GB. On the GPU: attention and DeltaNet mixers, router, shared experts, output head, MTP draft layer, KV cache, and an expert cache filling the remaining VRAM. In system RAM: all 24,576 routed experts pinned, with the CPU computing uncached experts in parallel with the GPU. The 28.8 GB n-gram table reads from SSD by default or from RAM.
Everything every token touches stays on the GPU. Spare VRAM caches the most-used experts; the CPU computes the rest straight from RAM.
  • On the GPU: the attention and DeltaNet mixers, the router, the shared experts, the output head, the MTP draft layer and the KV cache (with 128K context, only the most-read 32K stays here). Whatever VRAM is left becomes a cache of the most-used experts. The cache adapts as you chat.
  • In RAM: all 24,576 routed experts (48 × 512), pinned. When a token needs an expert that isn't cached, the CPU computes it in place, in parallel with the GPU, instead of copying it over.
  • The n-gram table: 28.8 GB. By default Strata reads it directly from SSD (--ple-io direct). I set --ple-io ram; I measured no speed difference.
  • MTP: the model ships with its own multi-token prediction layer. It drafts a few tokens ahead and the main model checks them in one pass, so the output is identical to running without it.

What you need: a 22 GB card, 64 GB of RAM at minimum, about 90 GB of disk

  • GPU: I used a modded 2080 Ti 22GB, and Strata used 21.6 GB of it. Strata compiled cleanly for sm_75 (Turing).
  • RAM: Strata's setup lists IQ3_S at 62 GB and says it "needs a 64 GB PC with little else running". On my box, available memory dropped by 47.6, 37.0 and 39.5 GiB across three loads, with the n-gram table in RAM. That's the whole machine's available memory, not Strata's own footprint.
  • Disk: the two GGUF shards are 54.8 GB + 28.8 GB = 83.6 GB. Add the MTP weights setup fetches (6.5 GiB) and the pack (1.5 GiB), and du puts the whole directory at about 86 GiB. Leave 90 GB.
  • Toolchain: Python 3.10+, g++ and nvcc. No sudo needed if those are already there. I built against CUDA 12.4.

The model is ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF, quant IQ3_S.

Install: one setup.py run builds the CUDA kernels and fetches the weights

Clone the repo and pin the commit I used:

git clone https://github.com/Niko1221/Strata
cd Strata
git checkout 6f32ec07

Then create a venv in the repo and run setup. I kept every cache on the NVMe drive:

python3 -m venv .venv

export XDG_CONFIG_HOME=/mnt/nvme1t/strata/config
export PIP_CACHE_DIR=/mnt/nvme1t/strata/pip-cache
export TMPDIR=/mnt/nvme1t/strata/tmp
export HF_HOME=/mnt/nvme1t/strata/data/hf

.venv/bin/python setup.py --family qwen --model IQ3_S --gpu 0 --cuda 12 --build \
  --context 32768 --kv int8 --vision no \
  --data-dir /mnt/nvme1t/strata/data \
  --port 8096 --host 127.0.0.1 --no-browser --no-start \
  --low-ram off --yes

A few notes:

  • --build --cuda 12 compiles from source. cuobjdump on the result shows 58 sm_75 ELF sections, so the 2080 Ti gets native kernels.
  • --low-ram off was needed because the 27B and ComfyUI were running during install, so setup saw too little free RAM.
  • Setup fetches the MTP weights from Qwen's official repo and checks each tensor's SHA-256.
  • I pre-downloaded the two shards with curl -C - (resumable) and checked their SHA-256 against Hugging Face.
  • It's done when setup prints All set.

Config, launch, first request: an OpenAI-compatible server on port 8096

My config file, strata-iq3_s-gpu2.json, sets these flags:

flagvaluewhat it does
--expert-cache / --prefillautoexpert cache size and prefill mode chosen by Strata
--spec / --spec-min-p4 / 0.5speculative drafting settings for the MTP layer
--max-context / --kv131072 / int8128K context, int8 KV cache
--kv-resident32768keeps the most-read 32K of KV on the GPU; the rest streams from RAM
--ple-ioramn-gram table in RAM instead of SSD reads

It also sets --mtp to the MTP weights directory, "parallel": 1 and the alias shio. Keep parallel at 1; the deep dive explains why.

To pin Strata to the third card, select it by UUID. Strata still sees it as GPU 0:

export CUDA_VISIBLE_DEVICES=<GPU UUID>
export STRATA_BF16_TC=1
.venv/bin/python serve/server.py --engine strata \
  --config strata-iq3_s-gpu2.json \
  --host 0.0.0.0 --port 8096 --gpu 0

Loading takes a few minutes. It's ready when /v1/models returns 200. Then:

curl -s http://localhost:8096/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "shio",
    "messages": [{"role": "user", "content": "Write a Python function that merges two sorted lists."}],
    "reasoning_effort": "low",
    "stream": false
  }'

Keep "reasoning_effort": "low". Strata turns thinking on at effort xhigh by default, which on a 140-question set took 3.4× as long for a lower score. Details are in the deep dive.

Running next to the 27B costs Strata 1.3% and the 27B 7.7%

The two models are on different GPUs, so the only shared resources are the CPU and memory bandwidth. I sent the same code prompt to both at once, 3 rounds.

aloneboth at oncechange
Strata (one card)70.3 tok/s69.4−1.3%
27B (two cards)243.1 tok/s224.4−7.7%

I'm assuming the contention is CPU or memory bandwidth; I didn't isolate it. Also, the overlap is short: the 27B finishes in about 3 s and Strata takes 10–12 s, so they only compete during Strata's first ~3 s.

The 27B never restarted. A health check every 5 s returned 106/106 HTTP 200.

The 243.1 uses this article's method (streaming, thinking off, 650 tokens). Part 22's 151.7 used a different method; don't compare them.

Sharing the card with ComfyUI: 5 minutes to switch to Shio, 12 seconds back

I call the Strata service on the third GPU Shio, after 汐 (tide). GPU2 runs Shio normally and switches to ComfyUI when I do video work.

I set Conflicts= between the two systemd user services, so starting one stops the other. A small script drives them:

~/bin/gpu2-mode shio     # switch to Strata
~/bin/gpu2-mode comfy    # switch to ComfyUI
~/bin/gpu2-mode status

The script has three guards:

  • It refuses to switch to Shio while ComfyUI's queue has jobs.
  • It waits for Shio to go idle before switching to ComfyUI.
  • If the target side isn't ready within 10 minutes, it rolls back to the previous side.

Measured over 3 rounds:

directionreadyfirst result
ComfyUI → Shio281–334 sfirst request 2.6–3.4 s after ready
Shio → ComfyUI12.4 sfirst 512px image 30.5–31 s (cold model load); second 18.3–18.4 s

From the command to the first ComfyUI image is about 43 s. Rounds 2 and 3 were faster than round 1, possibly from the OS page cache; I didn't isolate that. ComfyUI is still the boot default; Shio isn't enabled at boot.


Deep dive: parallel, thinking effort, 16K numbers, dead ends and versions

You can skip this section; everything you need to run the setup is above. This is the data behind the settings.

parallel: 2 slows even a single request

Strata's batch path doesn't use MTP drafting, but its docs say a request running alone still takes the MTP path. I measured a slowdown anyway, even with one request, and haven't tracked down why. Setting parallel back to 1 fixes it.

settingrequestsper requesttotal
parallel: 2160.32 tok/s54.49 (over full batch wall time)
parallel: 2230.39 / 32.2660.23
parallel: 1167.2, 70.1

Two concurrent requests together get less throughput than one request with parallel: 1. The ComfyUI switching test happened to run with parallel: 2 and saw 44.4, 60.3 and 44.8 tok/s. Back on parallel: 1, two checks gave 67.2 and 70.1.

Thinking defaults to xhigh; low scored higher in under a third of the time

Strata enables thinking at effort xhigh by default. I don't have a capability run on Shio itself, but I have the same 140-question set on the same model, Qwen3.8-Flash-Next, on a different machine and engine (GX10, vLLM, NVFP4):

effortscoretime
low1261,736 s
xhigh1175,902 s

Only the relative gap carries over.

Two ways to fix it:

  1. Send "reasoning_effort": "low" on every request.
  2. What I did: 3 lines in serve/server.py's OpenAI entry point. If a request carries no effort at all, fill in low.

Strata's /settings page has a shared default, but it only applies when the request has no thinking parameters at all. Clients that send only an "enable thinking" flag skip it. My client, Pi, does exactly that, so the server patch was the only thing that worked for it.

The 16K numbers were measured at 32K max context, one shot each

The 16K-prompt rows (TTFT, prefill, decode) ran with --max-context 32768, one shot per setting. Production uses 128K with --kv-resident 32768. I re-measured code on the production config: 70.1 tok/s, in line with the 70.3 from the main table.

Dual-card speed tricks that didn't help

I tried several things on the two-card setup; the next post covers them in detail.

  • --ple-io ram vs direct: no difference.
  • --peer-device (use NVLink between the cards): about 10% slower.
  • Using the third card as an expert helper for the two-card setup: code went from 91.4 to 48 tok/s (−47%).

The one flag that did help was STRATA_BF16_TC=1. On two cards it cut 16K TTFT from 12.6 s to 10.6 s. Single-card Shio uses that and nothing else from this list.

Versions

  • Strata engine 0.1.39, commit 6f32ec07, built with --build --cuda 12
  • CUDA 12.4, sm_75; cuobjdump shows 58 sm_75 ELF sections
  • Model: ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF, IQ3_S, two shards (54.8 GB + 28.8 GB)
  • MTP weights from Qwen's official repo, per-tensor SHA-256 checked
  • Draft vocab: Strata's built-in CJK subset, 106,299 ids
  • Env: STRATA_BF16_TC=1

What I didn't measure

  • The 140-question set on Shio itself.

Also in this series:

FAQ

Can a single RTX 2080 Ti run the 177B Qwen3.8-Flash-Next?
Yes, with Strata and enough system RAM. One modded 2080 Ti 22GB with 128 GB of RAM runs the IQ3_S quant at 70.3 tok/s on code, 52.9 on Chinese and 51.5 on English (streaming, thinking off, median of 3). The card uses 21.6 GB of VRAM. Two cards run the same prompts at 91.4, 60.0 and 60.0.
How does Strata fit a 177B MoE model on a 22 GB GPU?
The GPU holds the parts every token uses: attention and DeltaNet mixers, the router, shared experts, the output head, the MTP draft layer and the KV cache (at 128K context, only the most-read 32K of it). The rest of VRAM becomes a cache of the most-used experts, which adapts as you chat. All 24,576 routed experts sit pinned in system RAM, and the CPU computes any expert that is not cached, in parallel with the GPU.
How much system RAM does Strata need for Qwen3.8-Flash-Next IQ3_S?
Strata's setup lists IQ3_S at 62 GB and says it needs a 64 GB PC with little else running. On my 128 GB box, available system memory dropped by 37-48 GB across three loads, with the 28.8 GB n-gram table loaded into RAM.
Why is Strata slow or verbose on simple prompts?
Strata turns thinking on at effort xhigh by default. Send reasoning_effort low with each request, or patch the server to fill in low when a request carries no effort. On a 140-question set with the same model on other hardware, low scored 126 in 1,736 s and xhigh scored 117 in 5,902 s.
Can Strata and another model share the same machine?
On different GPUs, yes. With Strata on one 2080 Ti and a 27B on the other two, sending the same code prompt to both at once slowed Strata from 70.3 to 69.4 tok/s and the 27B from 243.1 to 224.4. The 27B finishes in about 3 s, so the overlap only covers the start of Strata's answer.

Read next

Don't miss the next one

Subscribe, and you won't.

One-click unsubscribe anytime.