~/ai-muninn
search
⌘K
blog
model anatomy
github
中
~ / blog
/
tag / tensor-parallel
❯
grep -r "#tensor-parallel" ~/blog
4 matches
date
read
title
2026-09-12
19m
[vLLM] Qwen3.8-Flash-Next TP=2 on Two DGX Sparks: 51.9 tok/s
#dgx-spark
#vllm
#tensor-parallel
#qwen3.8-flash-next
2026-08-23
20m
[Benchmark] Fixing NCCL's Stub-Library Error Cut PCIe Traffic 99% and Barely Moved Speed
#nccl
#llama.cpp
#tensor-parallel
#2080-ti
2026-08-23
23m
[Benchmark] Dual-GPU AllReduce Only Uses 8% of PCIe Gen3 x16 — It Moves Little, Very Often
#llama.cpp
#tensor-parallel
#pcie
#nvlink
2026-08-22
15m
[Benchmark] Two Modded 2080 Tis Reach 59.6 tok/s on Qwen3.8-27B With llama.cpp Tensor Parallel
#llama.cpp
#tensor-parallel
#2080-ti
#qwen3.8
← back to all posts