Measured on the machine, in the open.

Every number comes from a run we can repeat. Same 18-page archive, 12 questions, 6 pages of context. Each answer is scored on two things: the fact is correct, and the citation points to the right page. Raw results and scripts are in the bench folder.

Best model for each machine

Measured rows are ours. Projected rows scale by memory bandwidth until we test the machine. A nightly job refreshes the M2 and Jetson rows.

MachineChat modelEngineGen tok/sAnswer sStatus

Answer bench

Facts and Cites are counts out of 12. Median s is the answer time per question.

MachineModelFactsCitesMedian sTokensGen tok/s

LoCoMo

Public long-conversation memory benchmark. 304 questions from 2 conversations, gemma4:e4b on an M2, offline, k=6, token-F1. Reference points from the paper: human 87.9, GPT-3.5 with RAG 31.7 overall.

CategoryToken-F1
Multi-hop29.2
Temporal12.7
Open-domain17.8
Single-hop52.9
Adversarial53.5
Overall39.9

Small sample, a 4B model, no reranker, no date-aware retrieval. Temporal is the weak category for that last reason.

Embedding bench

Easy and Hard count the questions where the right page is in the top results. @6 counts hits with 6 pages of context. ms is the median embed time on the M2. nomic-embed-text is the default.

ModelDimsEasyHard@6ms
nomic-embed-text76811/129/1212/1213
embeddinggemma76810/128/1212/1231
bge-m3102410/129/1212/1243
mxbai-embed-large102410/129/1212/1223
qwen3-embedding 0.6B102410/1210/1211/1241
qwen3-embedding 4B256010/1210/1212/12144