Measured on the machine, in the open.
Every number comes from a run we can repeat. Same 18-page archive, 12 questions, 6 pages of context. Each answer is scored on two things: the fact is correct, and the citation points to the right page. Raw results and scripts are in the bench folder.
Best model for each machine
Measured rows are ours. Projected rows scale by memory bandwidth until we test the machine. A nightly job refreshes the M2 and Jetson rows.
| Machine | Chat model | Engine | Gen tok/s | Answer s | Status |
|---|
Answer bench
Facts and Cites are counts out of 12. Median s is the answer time per question.
| Machine | Model | Facts | Cites | Median s | Tokens | Gen tok/s |
|---|
LoCoMo
Public long-conversation memory benchmark. 304 questions from 2 conversations, gemma4:e4b on an M2, offline, k=6, token-F1. Reference points from the paper: human 87.9, GPT-3.5 with RAG 31.7 overall.
| Category | Token-F1 |
|---|---|
| Multi-hop | 29.2 |
| Temporal | 12.7 |
| Open-domain | 17.8 |
| Single-hop | 52.9 |
| Adversarial | 53.5 |
| Overall | 39.9 |
Small sample, a 4B model, no reranker, no date-aware retrieval. Temporal is the weak category for that last reason.
Embedding bench
Easy and Hard count the questions where the right page is in the top results. @6 counts hits with 6 pages of context. ms is the median embed time on the M2. nomic-embed-text is the default.
| Model | Dims | Easy | Hard | @6 | ms |
|---|---|---|---|---|---|
| nomic-embed-text | 768 | 11/12 | 9/12 | 12/12 | 13 |
| embeddinggemma | 768 | 10/12 | 8/12 | 12/12 | 31 |
| bge-m3 | 1024 | 10/12 | 9/12 | 12/12 | 43 |
| mxbai-embed-large | 1024 | 10/12 | 9/12 | 12/12 | 23 |
| qwen3-embedding 0.6B | 1024 | 10/12 | 10/12 | 11/12 | 41 |
| qwen3-embedding 4B | 2560 | 10/12 | 10/12 | 12/12 | 144 |