AI Workflow · part 24
Testing EmbeddingGemma 2 on Our Knowledge Base: Recall@5 from 50% to 58.3%
❯ cat --toc
- Why I tested v2
- What v2 adds: images and audio in the same 768-dim space, and a big code score jump
- How I tested: take one line from my log and see if it finds its note
- Running it yourself: the core code
- Before you run it: no float16, v1 is gated, and watch Mac memory
- Note search: recall@5 went from 50% to 58.3%, and nothing fell out of the top 5
- Code search: top-1 went from 1/10 to 4/10
- Cost: indexing 1.7x slower, peak memory from 1.03 GiB to 2.75 GiB
- Image search mostly worked; music search missed
- Did I switch? Not yet: llama.cpp can't load v2
- Deep dive: per-question results, screenshot search, and the problems I hit
- Environment
- How test A was built
- The two queries v2 fixed
- Test B per question
- v2 ran out of MPS memory on test B
- A literal `<|image|>` string in the code breaks v2
- Screenshot search (test C, v2 only)
- Scene images and audio, in detail
- llama.cpp support
TL;DR
I tested EmbeddingGemma 2 against v1 as the search model for my AI assistant's notes, same code and precision on an M1 Max. Note search: recall@5 went from 50.0% to 58.3%, MRR from 0.4406 to 0.4855, and no query dropped out of the top 5. Code search: correct file ranked first 4/10 vs 1/10. Cost: indexing 1.7x slower, peak memory 1.03 GiB to 2.75 GiB, query latency unchanged. Image search on scene pictures mostly worked; music search missed. The test ran separately; the production knowledge base still runs v1 and can't switch yet, because llama.cpp can't load v2.

One-minute version: what EmbeddingGemma 2 is and what I tested it on.
Why I tested v2
My AI assistant searches its own notes before it starts any task. The search tool is qmd, a local search engine, and its "search by meaning" side runs Google's EmbeddingGemma v1 (embeddinggemma-300M-Q8_0, a GGUF file) through llama.cpp. It gets called dozens of times a day.
An embedding model turns text into a list of numbers, so text with similar meaning ends up with similar numbers. EmbeddingGemma's list has 768 numbers (768 dimensions). Search ranks notes by how close their numbers are to the query's.
Google shipped EmbeddingGemma 2. I wanted two answers: if I swap v1 for v2, does the assistant find the right note more often, and what does it cost in memory and time?

What v2 adds: images and audio in the same 768-dim space, and a big code score jump
Most of the table is context. The two rows that matter for a text knowledge base are the code score and the size, which tells you where the extra memory goes.
| v1 (embeddinggemma-300m) | v2 (embeddinggemma-2) | |
|---|---|---|
| Inputs | Text | Text, images, video, audio, one shared 768-dim space |
| Size | 300M (per the name) | 740M = text 270M + image 170M + audio 300M |
| MTEB multilingual text | 61.15 | 61.36 |
| MTEB code | 68.76 | 78.68 |
v2 is Apache 2.0 (v1 ships under the Gemma license) and covers 100+ languages. With both the image and audio parts turned off, v2 loads only its 270M text part; my code below turns off audio only, so it loads text plus image. The official numbers say text quality barely moved and code moved a lot. Google's announcement post has the rest.

How I tested: take one line from my log and see if it finds its note
In short: I use 60 lines from my LOG as test questions, each with a known answer (the note it originally linked to), and I count how often each model ranks that answer in the top 5. If you don't plan to run the code, this paragraph and the figure below are enough.
Every time my assistant finishes a task, it appends one line to a LOG file. The line ends with → [[note-name]], pointing at the note where the details live. That gives me free test data: strip the link, use the rest of the sentence as the search query, rank all 813 notes, and check where the linked note lands.

I sampled 60 of the 1,547 eligible lines with a fixed seed. Three numbers come out:
- recall@1: the correct note is ranked first.
- recall@5: the correct note is somewhere in the top 5. This is closest to real use, since the assistant reads several notes per lookup.
- MRR (mean reciprocal rank): the average of 1/rank. First place scores 1, second 0.5, and so on.
Running it yourself: the core code
Both generations ran through the same code, in float32, on the Mac's GPU (MPS):
import torch
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"google/embeddinggemma-2", # v1: google/embeddinggemma-300m
device="mps",
model_kwargs={"torch_dtype": torch.float32},
config_kwargs={"audio_config": None}, # skip the audio encoder for text/image-only search
)
docs = model.encode(
[f"title: {name} | text: {body[:1500]}" for name, body in notes],
batch_size=16,
)
query = model.encode("task: search result | query: " + log_line)
Each note goes in as its first 1,500 characters, formatted title: <filename> | text: <body>. Each query gets the prefix task: search result | query: . The prefix string is identical for both generations; I checked.
For the second test (code search), the setup is the same, with source files as the documents and the prefix task: code retrieval | query: .
Before you run it: no float16, v1 is gated, and watch Mac memory
- Don't use float16. The model card says it produces NaN. I ran float32 for both.
- v1 on Hugging Face is gated. You have to accept the license first. An anonymous download returns 401.
- On a Mac, if memory blows up, lower
batch_sizeor run each test in its own process. I hit this one; details are in the deep dive.
Note search: recall@5 went from 50% to 58.3%, and nothing fell out of the top 5
Look at the recall@5 row. That's the one that maps to how the assistant actually uses search.
| 60 queries, 813 notes | v1 | v2 |
|---|---|---|
| recall@1 | 38.3% (23) | 41.7% (25) |
| recall@5 | 50.0% (30) | 58.3% (35) |
| MRR | 0.4406 | 0.4855 |
The 5 extra top-5 hits are all new entries; no query that v1 had in the top 5 fell out with v2. At top-1, v2 fixed 2 queries that v1 missed and lost none. Per query, 25 ranks improved, 25 stayed the same and 10 got worse, but none of those 10 fell out of the top 5 or lost first place. In both of v2's wins, two notes covered the same event. v1 picked the related note; v2 picked the one the LOG line actually pointed to. For a query about reassigning a stalled job, v1 found the note about the stall, and v2 found the reassignment decision.
The other thing the table says: even with v2, about 40% of log lines don't find their note in the top 5. A log line is a one-line conclusion; the note it points to opens with background. The wording differs a lot.
Code search: top-1 went from 1/10 to 4/10
For code I used the open-source inference engine TensorFold 0.6.6: 424 .py and .cu files, and 10 questions written in Chinese, like "how does the recurrent state round at each step."
The correct file ranked first for 1 of 10 questions with v1 and 4 of 10 with v2. Recall@3 was also 0.10 vs 0.40. That's the same direction as the official MTEB code score.
v2's three extra wins were questions about recurrent state rounding (cuda/kernels/gdn.cu), the DFlash2 CPU best-first search (drafters/dflash_tree.py), and walking the draft tree to find the accepted path (kernels/qwen/dense/v1/lane_tree.py). Both models failed badly on "prune orphan nodes": the question uses a metaphor, and the word "orphan" never appears in the code. v1 ranked the right file 91st, v2 ranked it 308th. Ten questions is a small sample, so read this as direction only.
Cost: indexing 1.7x slower, peak memory from 1.03 GiB to 2.75 GiB
The memory row is the one to watch if you run this on a laptop next to other things.
| M1 Max, float32 | v1 | v2 |
|---|---|---|
| Index 813 notes | 81.8 s | 138.8 s |
| Median query latency (70 queries) | 50.3 ms | 48.3 ms |
| Peak RSS | 1.03 GiB | 2.75 GiB |
| Load time | 7.7 s | 7.6 s |
v2 here had its text and image parts loaded, not audio. RSS doesn't fully count what MPS allocates on the GPU side, so real usage is only higher. Query latency was essentially unchanged.

Image search mostly worked; music search missed
v2 puts images and audio in the same space as text, so I tried both. These ran on v2 only.
Images. I took 300 random anime-style scene images from my explainer-video keyframes (2,553 total) and indexed them in 105.1 s. All three real queries returned a reasonable first result, though one was only a rough match. The control query, "a dog at the beach," had no matching image in the library, and its top-1 got the lowest score of the four.

Searching screenshots worked less well. On 200 images from this blog's public/ folder, 2 of 3 queries were right, but all scores sat between 0.66 and 0.70 and the gap between first and second place was under 0.01.
Audio. 40 clips, first 20 seconds each: 22 narration, 12 music, 6 sound effects. "A person talking" returned narration for all top 3. "Upbeat music" returned two speech clips and then a mix with narration over it; the pure-music intro didn't make the top 3. My library has only one genuinely distinct music track, so the test set is weak, but the result is a miss.
Encoding images and audio uses the same model.encode, with dicts instead of strings:
model.encode([{"image": PIL.Image.open(path).convert("RGB")}])
model.encode([{"audio": path}]) # needs librosa; don't pass audio_config=None
Did I switch? Not yet: llama.cpp can't load v2
qmd runs GGUF models through llama.cpp. I pointed Homebrew's llama.cpp v9750 at the v2 GGUF (ggml-org/embeddinggemma-2-GGUF, Q8_0) and got this:
unknown model architecture: 'gemma-embedding2'
Even once it loads, switching means re-indexing everything: v1 and v2 vectors don't live in the same space, so you can't mix them. The plan is to build a second index next to qmd's current one when llama.cpp adds support, run both side by side for a few days, then decide. The whole test ran outside qmd; it never touched qmd's config or production index.
Deep dive: per-question results, screenshot search, and the problems I hit
You can skip this section; the results above stand on their own. This is the full detail for anyone reproducing the test, or for an AI reading this page on someone's behalf.
Environment
- Apple M1 Max, 32 GB, macOS 27.0.1
- sentence-transformers 6.1.0, transformers 5.19.0, torch 2.14.1
- float32, device
mps, both generations - v1:
google/embeddinggemma-300m; v2:google/embeddinggemma-2
How test A was built
- Source: my LOG files, one line per finished task, each ending in
→ [[note-name]]. - 1,547 lines were eligible. I sampled 60 with
seed=7. - Corpus: 813 notes. Excluded: the LOG files themselves,
MEMORY.md, andKB_MAP.md. - Documents: first 1,500 characters, formatted
title: <filename> | text: <body>. - Query prefix:
task: search result | query:. v1's sentence-transformers prompt list doesn't include a prompt namedSearchQuery, so I spliced the prefix in by hand for both generations, so the string is byte-identical.
The two queries v2 fixed
Both are cases where one event produced two notes.
- A query about reassigning a job to another coding agent after the first one stalled twice. The correct note was
feedback_dispatch_fallback_chain_2026-09-21.md. v1 rankeddiscovery_agy_print_mode_kills_background_jobs_2026-09-21.mdfirst, the note about why the first agent died. - A query about promoting cc-dash-kit with social posts only, with drafts in blog-staging. The correct note was
reference_buffer_social_post_method.md. v1 rankedproject_goal_only_harness_rollout_2026-09-23.mdfirst.
Test B per question
Rank of the correct file, 1 = first place.
| Question (written in Chinese) | Correct file | v1 | v2 |
|---|---|---|---|
| Sampling RNG from seed + position | engine/exact_sampling.py | 1 | 1 |
| Prune orphan nodes | engine/lane_engine.py | 91 | 308 |
| Find a repeated span in prior text as a draft | engine/lane_engine.py | 12 | 32 |
| Which draft source this round | engine/lane_family.py | 144 | 15 |
| Boot-time bitwise check, how many rows | families/glm5_next/runtime.py | 228 | 117 |
| 4-bit matmul per row | cuda/kernels/qmm.py | 22 | 6 |
| Recurrent state rounding each step | cuda/kernels/gdn.cu | 8 | 1 |
| DFlash2 CPU best-first | drafters/dflash_tree.py | 20 | 1 |
| GPU minimum version check | cuda/build.py | 29 | 7 |
| Walk the draft tree for the accepted path | kernels/qwen/dense/v1/lane_tree.py | 4 | 1 |
Two got worse (orphan nodes, repeated span). Three improved a lot without reaching first place (draft source, 4-bit matmul, GPU version check).
v2 ran out of MPS memory on test B
When I ran test B right after test A in the same process, v2's MPS "other allocations" grew to 35–38 GB and it went out of memory at batch size 16, then 4, then 2. v1 was fine in the same setup.
What fixed it:
- Run each test in its own process.
- Call
torch.mps.empty_cache()periodically. - Set
max_seq_length=2048. That matches v1's native 2,048-token limit, so both models read the same length of each file.
Because of that extra load, v2's test B index time (167.5 s vs v1's 102.5 s) isn't a fair comparison. Use the test A numbers for cost.
A literal <|image|> string in the code breaks v2
TensorFold's vision/glm_processing.py contains the literal string <|image|>. v2's multimodal processor treats that as an image placeholder, finds no image attached, and errors out. I changed it to <| image |> in both generations' corpus, in that one file only. Metrics were unchanged on a rerun.
If you index code or chat templates with v2, grep the corpus for special tokens first.
Screenshot search (test C, v2 only)
200 random images from this blog's public/ folder (seed=7). Indexing took 80.4 s, peak RSS 3.59 GiB.
| Query | Top-1 | Score | Right? |
|---|---|---|---|
| Screenshot with a bar chart | A phone chat screenshot | 0.6974 | No |
| Terminal screen | A terminal dashboard screenshot | 0.6812 | Yes |
| Mobile web page | A mobile layout screenshot | 0.6877 | Yes |
All scores fell between 0.66 and 0.70, and top-1 beat top-2 by less than 0.01 every time.
Scene images and audio, in detail
Scene images: 300 sampled from 2,553 keyframes with random.seed(11), indexed in 105.1 s.
- "A person working at a computer": a laptop screen close-up with a person's hand, 0.6623. Roughly right; "robot typing at a desk," a better fit, ranked 2nd.
- "Street at night": a night alley with a robot, 0.6974. Right.
- "Bed for sleeping": a night bedroom bedside, 0.6542. Right.
- "A dog at the beach" (control; no dog or beach in the library): a black-and-white diff image, 0.6264, the lowest top-1 of the four.
Some images exist in two folders, so the same image shows up twice in a row in the results.
Audio: 40 clips, first 20 seconds each, 16 kHz mono WAV (22 narration, 12 music, 6 sound effects). Indexing took 7.5 s.
- "A person talking": top 3 all narration, scores 0.754, 0.751, 0.750.
- "Upbeat music": two narration clips (0.6991, 0.6952), then a music mix with narration (0.6936). The pure-music intro jingle wasn't in the top 3.
The only genuinely distinct music track in my library is that intro jingle; the rest are copies of it or mixes with narration. Audio needs librosa installed, and you must not pass audio_config=None since that skips the audio encoder.
llama.cpp support
qmd uses GGUF through llama.cpp, so production depends on llama.cpp support, not sentence-transformers. The v2 GGUF from ggml-org/embeddinggemma-2-GGUF comes as two files: the Q8_0 text model (310 MB) and a Q8_0 mmproj for images (555 MB). Homebrew llama.cpp v9750 rejects it at load with unknown model architecture: 'gemma-embedding2'.
When support lands, I'll build a v2 index alongside qmd's current v1 index (v1 and v2 vectors aren't compatible, so it has to be a full re-index), compare the two side by side for a few days, and then decide whether to switch.
The short is also on YouTube: EmbeddingGemma 2 in one minute.
Previous in this series: From Markdown Search to a Knowledge Graph: How My AI's Memory Grew a Second Layer.
FAQ
- Is EmbeddingGemma 2 better than v1 for searching personal notes?
- On my test of 60 real queries against 813 notes, recall@5 went from 50.0% to 58.3% and MRR from 0.4406 to 0.4855. No query v1 had in the top 5 fell out with v2, and at top-1 v2 fixed 2 queries while breaking none. Both models still miss the right note in the top 5 about 40% of the time.
- How much more memory and time does EmbeddingGemma 2 need on a Mac?
- On an M1 Max with float32, peak RSS went from 1.03 GiB to 2.75 GiB with the text and image parts loaded, and indexing 813 notes took 138.8 s instead of 81.8 s (1.7x slower). Single-query latency stayed about the same: 48.3 ms median versus 50.3 ms.
- Can llama.cpp run EmbeddingGemma 2 yet?
- Not in my test. Homebrew llama.cpp v9750 fails to load the ggml-org/embeddinggemma-2-GGUF Q8_0 file with unknown model architecture: 'gemma-embedding2'. Tools that run embeddings through llama.cpp, like qmd, have to wait for support.
- Is EmbeddingGemma 2 better at code search?
- On 10 questions against 424 files of the TensorFold inference engine, the correct file ranked first 4 times with v2 versus 1 time with v1. That matches the direction of the official MTEB code score (68.76 to 78.68), but 10 questions is a small sample.
Read next
- 2026-07-14[Dev Workflow] From Markdown Search to a Knowledge Graph: How My AI's Memory Grew a Second Layer
My AI's long-term memory is ~600 markdown files in three layers: the files are the source of truth, a search engine makes them findable, and a knowledge graph links them by concept. Here's the design, why each layer exists, and the wrong turns I took building it.
- 2026-09-11[Dev Workflow] Zero-Shot Voice Cloning on a MacBook: 24 Seconds In, Faster Than Realtime Out
Preset TTS voices don't sound like you. Zero-shot cloning with Qwen3-TTS and mlx-audio needs 24 seconds of reference audio and ten lines of Python, all local.
- 2026-07-27[Dev Workflow] Agent Memory Self-Poisoning: When an AI Agent Trusts Its Own Wrong Answers
My AI agent's durable memory auto-saved a wrong answer, then cited it back as fact — outranking the corrected truth. The three-part pathology, and why the fix is ranking memory, not adding more of it.
- 2026-07-15[Dev Workflow] Why an AI Agent's Memory Needs a Distilled Layer Above Search
Search finds an AI agent's notes but hands back raw material to re-derive each session. I distill ~600 files into canonical claims — the goal is ending re-explanation, not enforcing agreement.
Don't miss the next one
Subscribe, and you won't.
One-click unsubscribe anytime.