~/blog/embeddinggemma-2-knowledge-base-test

AI Workflow · part 24

Testing EmbeddingGemma 2 on Our Knowledge Base: Recall@5 from 50% to 58.3%

❯ cat --toc

TL;DR

I tested EmbeddingGemma 2 against v1 as the search model for my AI assistant's notes, same code and precision on an M1 Max. Note search: recall@5 went from 50.0% to 58.3%, MRR from 0.4406 to 0.4855, and no query dropped out of the top 5. Code search: correct file ranked first 4/10 vs 1/10. Cost: indexing 1.7x slower, peak memory 1.03 GiB to 2.75 GiB, query latency unchanged. Image search on scene pictures mostly worked; music search missed. The test ran separately; the production knowledge base still runs v1 and can't switch yet, because llama.cpp can't load v2.

EmbeddingGemma 2 test card: note search recall@5 from 50% to 58.3%, code search top-1 from 1/10 to 4/10, memory from 1.03 to 2.75 GiB, also searches images and audio

One-minute version: what EmbeddingGemma 2 is and what I tested it on.


Why I tested v2

My AI assistant searches its own notes before it starts any task. The search tool is qmd, a local search engine, and its "search by meaning" side runs Google's EmbeddingGemma v1 (embeddinggemma-300M-Q8_0, a GGUF file) through llama.cpp. It gets called dozens of times a day.

An embedding model turns text into a list of numbers, so text with similar meaning ends up with similar numbers. EmbeddingGemma's list has 768 numbers (768 dimensions). Search ranks notes by how close their numbers are to the query's.

Google shipped EmbeddingGemma 2. I wanted two answers: if I swap v1 for v2, does the assistant find the right note more often, and what does it cost in memory and time?

Google Gemma post on X announcing EmbeddingGemma 2: a lightweight multimodal embedding model that maps text, code, images, video and audio into one space, with a 740M parameter form factor
The launch post from @googlegemma on X (source: x.com/googlegemma)

What v2 adds: images and audio in the same 768-dim space, and a big code score jump

Most of the table is context. The two rows that matter for a text knowledge base are the code score and the size, which tells you where the extra memory goes.

v1 (embeddinggemma-300m)v2 (embeddinggemma-2)
InputsTextText, images, video, audio, one shared 768-dim space
Size300M (per the name)740M = text 270M + image 170M + audio 300M
MTEB multilingual text61.1561.36
MTEB code68.7678.68

v2 is Apache 2.0 (v1 ships under the Gemma license) and covers 100+ languages. With both the image and audio parts turned off, v2 loads only its 270M text part; my code below turns off audio only, so it loads text plus image. The official numbers say text quality barely moved and code moved a lot. Google's announcement post has the rest.

Text rows of the official model card benchmark table: MTEB multilingual v2 61.36 vs v1 61.15; MTEB code v2 78.68 vs v1 68.76
The two text rows in the model card's benchmark table; the image, video and audio rows only have v2 scores and are cropped out (source: Hugging Face, google/embeddinggemma-2)

How I tested: take one line from my log and see if it finds its note

In short: I use 60 lines from my LOG as test questions, each with a known answer (the note it originally linked to), and I count how often each model ranks that answer in the top 5. If you don't plan to run the code, this paragraph and the figure below are enough.

Every time my assistant finishes a task, it appends one line to a LOG file. The line ends with → [[note-name]], pointing at the note where the details live. That gives me free test data: strip the link, use the rest of the sentence as the search query, rank all 813 notes, and check where the linked note lands.

Test flow: one LOG line, strip the link, use it as the query, rank all 813 notes, check whether the linked note is in the top 5

I sampled 60 of the 1,547 eligible lines with a fixed seed. Three numbers come out:

  • recall@1: the correct note is ranked first.
  • recall@5: the correct note is somewhere in the top 5. This is closest to real use, since the assistant reads several notes per lookup.
  • MRR (mean reciprocal rank): the average of 1/rank. First place scores 1, second 0.5, and so on.

Running it yourself: the core code

Both generations ran through the same code, in float32, on the Mac's GPU (MPS):

import torch
from sentence_transformers import SentenceTransformer

model = SentenceTransformer(
    "google/embeddinggemma-2",          # v1: google/embeddinggemma-300m
    device="mps",
    model_kwargs={"torch_dtype": torch.float32},
    config_kwargs={"audio_config": None},  # skip the audio encoder for text/image-only search
)

docs = model.encode(
    [f"title: {name} | text: {body[:1500]}" for name, body in notes],
    batch_size=16,
)
query = model.encode("task: search result | query: " + log_line)

Each note goes in as its first 1,500 characters, formatted title: <filename> | text: <body>. Each query gets the prefix task: search result | query: . The prefix string is identical for both generations; I checked.

For the second test (code search), the setup is the same, with source files as the documents and the prefix task: code retrieval | query: .

Before you run it: no float16, v1 is gated, and watch Mac memory

  • Don't use float16. The model card says it produces NaN. I ran float32 for both.
  • v1 on Hugging Face is gated. You have to accept the license first. An anonymous download returns 401.
  • On a Mac, if memory blows up, lower batch_size or run each test in its own process. I hit this one; details are in the deep dive.

Note search: recall@5 went from 50% to 58.3%, and nothing fell out of the top 5

Look at the recall@5 row. That's the one that maps to how the assistant actually uses search.

60 queries, 813 notesv1v2
recall@138.3% (23)41.7% (25)
recall@550.0% (30)58.3% (35)
MRR0.44060.4855

The 5 extra top-5 hits are all new entries; no query that v1 had in the top 5 fell out with v2. At top-1, v2 fixed 2 queries that v1 missed and lost none. Per query, 25 ranks improved, 25 stayed the same and 10 got worse, but none of those 10 fell out of the top 5 or lost first place. In both of v2's wins, two notes covered the same event. v1 picked the related note; v2 picked the one the LOG line actually pointed to. For a query about reassigning a stalled job, v1 found the note about the stall, and v2 found the reassignment decision.

The other thing the table says: even with v2, about 40% of log lines don't find their note in the top 5. A log line is a one-line conclusion; the note it points to opens with background. The wording differs a lot.

Code search: top-1 went from 1/10 to 4/10

For code I used the open-source inference engine TensorFold 0.6.6: 424 .py and .cu files, and 10 questions written in Chinese, like "how does the recurrent state round at each step."

The correct file ranked first for 1 of 10 questions with v1 and 4 of 10 with v2. Recall@3 was also 0.10 vs 0.40. That's the same direction as the official MTEB code score.

v2's three extra wins were questions about recurrent state rounding (cuda/kernels/gdn.cu), the DFlash2 CPU best-first search (drafters/dflash_tree.py), and walking the draft tree to find the accepted path (kernels/qwen/dense/v1/lane_tree.py). Both models failed badly on "prune orphan nodes": the question uses a metaphor, and the word "orphan" never appears in the code. v1 ranked the right file 91st, v2 ranked it 308th. Ten questions is a small sample, so read this as direction only.

Cost: indexing 1.7x slower, peak memory from 1.03 GiB to 2.75 GiB

The memory row is the one to watch if you run this on a laptop next to other things.

M1 Max, float32v1v2
Index 813 notes81.8 s138.8 s
Median query latency (70 queries)50.3 ms48.3 ms
Peak RSS1.03 GiB2.75 GiB
Load time7.7 s7.6 s

v2 here had its text and image parts loaded, not audio. RSS doesn't fully count what MPS allocates on the GPU side, so real usage is only higher. Query latency was essentially unchanged.

Bar chart: note search recall@5 v1 50.0% vs v2 58.3%; code search top-1 hits v1 1/10 vs v2 4/10; peak memory v1 1.03 GiB vs v2 2.75 GiB
All three measurements side by side; for memory, lower is better; the v2 bar is amber (M1 Max, float32)

Image search mostly worked; music search missed

v2 puts images and audio in the same space as text, so I tried both. These ran on v2 only.

Images. I took 300 random anime-style scene images from my explainer-video keyframes (2,553 total) and indexed them in 105.1 s. All three real queries returned a reasonable first result, though one was only a rough match. The control query, "a dog at the beach," had no matching image in the library, and its top-1 got the lowest score of the four.

v2 image search demo: four text queries (labels in the image are in Chinese) with their top results. A person working at a computer: a laptop screen close-up with a hand, score 0.6623. Street at night: a night alley with a robot, 0.6974. Bed for sleeping: a night bedroom bedside, 0.6542. A dog at the beach (control, no such image exists): a black-and-white diff image, 0.6264.

Searching screenshots worked less well. On 200 images from this blog's public/ folder, 2 of 3 queries were right, but all scores sat between 0.66 and 0.70 and the gap between first and second place was under 0.01.

Audio. 40 clips, first 20 seconds each: 22 narration, 12 music, 6 sound effects. "A person talking" returned narration for all top 3. "Upbeat music" returned two speech clips and then a mix with narration over it; the pure-music intro didn't make the top 3. My library has only one genuinely distinct music track, so the test set is weak, but the result is a miss.

Encoding images and audio uses the same model.encode, with dicts instead of strings:

model.encode([{"image": PIL.Image.open(path).convert("RGB")}])
model.encode([{"audio": path}])  # needs librosa; don't pass audio_config=None

Did I switch? Not yet: llama.cpp can't load v2

qmd runs GGUF models through llama.cpp. I pointed Homebrew's llama.cpp v9750 at the v2 GGUF (ggml-org/embeddinggemma-2-GGUF, Q8_0) and got this:

unknown model architecture: 'gemma-embedding2'

Even once it loads, switching means re-indexing everything: v1 and v2 vectors don't live in the same space, so you can't mix them. The plan is to build a second index next to qmd's current one when llama.cpp adds support, run both side by side for a few days, then decide. The whole test ran outside qmd; it never touched qmd's config or production index.

Deep dive: per-question results, screenshot search, and the problems I hit

You can skip this section; the results above stand on their own. This is the full detail for anyone reproducing the test, or for an AI reading this page on someone's behalf.

Environment

  • Apple M1 Max, 32 GB, macOS 27.0.1
  • sentence-transformers 6.1.0, transformers 5.19.0, torch 2.14.1
  • float32, device mps, both generations
  • v1: google/embeddinggemma-300m; v2: google/embeddinggemma-2

How test A was built

  • Source: my LOG files, one line per finished task, each ending in → [[note-name]].
  • 1,547 lines were eligible. I sampled 60 with seed=7.
  • Corpus: 813 notes. Excluded: the LOG files themselves, MEMORY.md, and KB_MAP.md.
  • Documents: first 1,500 characters, formatted title: <filename> | text: <body>.
  • Query prefix: task: search result | query: . v1's sentence-transformers prompt list doesn't include a prompt named SearchQuery, so I spliced the prefix in by hand for both generations, so the string is byte-identical.

The two queries v2 fixed

Both are cases where one event produced two notes.

  1. A query about reassigning a job to another coding agent after the first one stalled twice. The correct note was feedback_dispatch_fallback_chain_2026-09-21.md. v1 ranked discovery_agy_print_mode_kills_background_jobs_2026-09-21.md first, the note about why the first agent died.
  2. A query about promoting cc-dash-kit with social posts only, with drafts in blog-staging. The correct note was reference_buffer_social_post_method.md. v1 ranked project_goal_only_harness_rollout_2026-09-23.md first.

Test B per question

Rank of the correct file, 1 = first place.

Question (written in Chinese)Correct filev1v2
Sampling RNG from seed + positionengine/exact_sampling.py11
Prune orphan nodesengine/lane_engine.py91308
Find a repeated span in prior text as a draftengine/lane_engine.py1232
Which draft source this roundengine/lane_family.py14415
Boot-time bitwise check, how many rowsfamilies/glm5_next/runtime.py228117
4-bit matmul per rowcuda/kernels/qmm.py226
Recurrent state rounding each stepcuda/kernels/gdn.cu81
DFlash2 CPU best-firstdrafters/dflash_tree.py201
GPU minimum version checkcuda/build.py297
Walk the draft tree for the accepted pathkernels/qwen/dense/v1/lane_tree.py41

Two got worse (orphan nodes, repeated span). Three improved a lot without reaching first place (draft source, 4-bit matmul, GPU version check).

v2 ran out of MPS memory on test B

When I ran test B right after test A in the same process, v2's MPS "other allocations" grew to 35–38 GB and it went out of memory at batch size 16, then 4, then 2. v1 was fine in the same setup.

What fixed it:

  • Run each test in its own process.
  • Call torch.mps.empty_cache() periodically.
  • Set max_seq_length=2048. That matches v1's native 2,048-token limit, so both models read the same length of each file.

Because of that extra load, v2's test B index time (167.5 s vs v1's 102.5 s) isn't a fair comparison. Use the test A numbers for cost.

A literal <|image|> string in the code breaks v2

TensorFold's vision/glm_processing.py contains the literal string <|image|>. v2's multimodal processor treats that as an image placeholder, finds no image attached, and errors out. I changed it to <| image |> in both generations' corpus, in that one file only. Metrics were unchanged on a rerun.

If you index code or chat templates with v2, grep the corpus for special tokens first.

Screenshot search (test C, v2 only)

200 random images from this blog's public/ folder (seed=7). Indexing took 80.4 s, peak RSS 3.59 GiB.

QueryTop-1ScoreRight?
Screenshot with a bar chartA phone chat screenshot0.6974No
Terminal screenA terminal dashboard screenshot0.6812Yes
Mobile web pageA mobile layout screenshot0.6877Yes

All scores fell between 0.66 and 0.70, and top-1 beat top-2 by less than 0.01 every time.

Scene images and audio, in detail

Scene images: 300 sampled from 2,553 keyframes with random.seed(11), indexed in 105.1 s.

  • "A person working at a computer": a laptop screen close-up with a person's hand, 0.6623. Roughly right; "robot typing at a desk," a better fit, ranked 2nd.
  • "Street at night": a night alley with a robot, 0.6974. Right.
  • "Bed for sleeping": a night bedroom bedside, 0.6542. Right.
  • "A dog at the beach" (control; no dog or beach in the library): a black-and-white diff image, 0.6264, the lowest top-1 of the four.

Some images exist in two folders, so the same image shows up twice in a row in the results.

Audio: 40 clips, first 20 seconds each, 16 kHz mono WAV (22 narration, 12 music, 6 sound effects). Indexing took 7.5 s.

  • "A person talking": top 3 all narration, scores 0.754, 0.751, 0.750.
  • "Upbeat music": two narration clips (0.6991, 0.6952), then a music mix with narration (0.6936). The pure-music intro jingle wasn't in the top 3.

The only genuinely distinct music track in my library is that intro jingle; the rest are copies of it or mixes with narration. Audio needs librosa installed, and you must not pass audio_config=None since that skips the audio encoder.

llama.cpp support

qmd uses GGUF through llama.cpp, so production depends on llama.cpp support, not sentence-transformers. The v2 GGUF from ggml-org/embeddinggemma-2-GGUF comes as two files: the Q8_0 text model (310 MB) and a Q8_0 mmproj for images (555 MB). Homebrew llama.cpp v9750 rejects it at load with unknown model architecture: 'gemma-embedding2'.

When support lands, I'll build a v2 index alongside qmd's current v1 index (v1 and v2 vectors aren't compatible, so it has to be a full re-index), compare the two side by side for a few days, and then decide whether to switch.

The short is also on YouTube: EmbeddingGemma 2 in one minute.

Previous in this series: From Markdown Search to a Knowledge Graph: How My AI's Memory Grew a Second Layer.

FAQ

Is EmbeddingGemma 2 better than v1 for searching personal notes?
On my test of 60 real queries against 813 notes, recall@5 went from 50.0% to 58.3% and MRR from 0.4406 to 0.4855. No query v1 had in the top 5 fell out with v2, and at top-1 v2 fixed 2 queries while breaking none. Both models still miss the right note in the top 5 about 40% of the time.
How much more memory and time does EmbeddingGemma 2 need on a Mac?
On an M1 Max with float32, peak RSS went from 1.03 GiB to 2.75 GiB with the text and image parts loaded, and indexing 813 notes took 138.8 s instead of 81.8 s (1.7x slower). Single-query latency stayed about the same: 48.3 ms median versus 50.3 ms.
Can llama.cpp run EmbeddingGemma 2 yet?
Not in my test. Homebrew llama.cpp v9750 fails to load the ggml-org/embeddinggemma-2-GGUF Q8_0 file with unknown model architecture: 'gemma-embedding2'. Tools that run embeddings through llama.cpp, like qmd, have to wait for support.
Is EmbeddingGemma 2 better at code search?
On 10 questions against 424 files of the TensorFold inference engine, the correct file ranked first 4 times with v2 versus 1 time with v1. That matches the direction of the official MTEB code score (68.76 to 78.68), but 10 questions is a small sample.

Read next

Don't miss the next one

Subscribe, and you won't.

One-click unsubscribe anytime.