~/ai-muninn
search
⌘K
blog
github
中
~ / blog
/
tag / gguf
❯
grep -r "#gguf" ~/blog
4 matches
date
read
title
2026-08-20
11m
[Just for Fun — Advanced] Why Isn't Your 4-Bit Quant Faster on a 2080 Ti? I Tore Open the CUDA Backend to Find Out
#2080-ti
#turing
#quantization
#gguf
2026-08-18
7m
[Junk-Tier Big Models #5] Qwen3.8-27B on a 2018 22GB card: 30 tok/s, and it one-shot a 3D scene
#qwen3.8
#unsloth
#2080-ti
#llama.cpp
2026-07-23
13m
[LLM Deep Dive] Surgical GGUF Quantization: Quantize Only the Tensors You Choose
#quantization
#gguf
#llama.cpp
#llama-quantize
2026-04-10
12m
[LLM 101 #4] What Is Quantization? Q4, Q8, FP16 Explained
#llm
#quantization
#ollama
#beginner
← back to all posts