~/ai-muninn
search
⌘K
blog
github
中
~ / blog
/
tag / gtx-970
❯
grep -r "#gtx-970" ~/blog
4 matches
date
read
title
2026-06-14
14m
[Just for Fun] A blog RAG support bot on a GTX 970: no torch, no vector DB, no LangChain
#gemma-4
#gtx-970
#rag
#llama.cpp
2026-06-14
10m
[Just for Fun] On a GTX 970, Flash Attention nearly doubles long-context decode (24.3 → 42.5 tok/s)
#gemma-4
#gtx-970
#flash-attention
#kv-cache
2026-06-09
10m
[Just for Fun] A GTX 970 as an offline voice assistant: Gemma 4 E2B + Piper TTS (2.8s end-to-end)
#gemma-4
#gtx-970
#multimodal
#piper-tts
2026-06-09
10m
[Just for Fun] Gemma 4 E2B on a GTX 970: the biggest quant runs fastest (47.6 tok/s)
#gemma-4
#quantization
#gtx-970
#llama.cpp
← back to all posts