0verload · live memory benchmarks

Four 8k-context qwen-7B models, answering million-token benchmark questions — live.

Nobody would point a model with an 8,000-token window at a benchmark whose conversations run to a million tokens and beyond. Who would even try that? We run four of them at once — tiny qwen-7B bots, one per consumer GPU — and they pass anyway. Because the model isn’t doing the remembering: 0verload is. The model reasons over what it’s handed; the store hands it back the right fact in milliseconds. The number that matters is recall latency — and every question, answer, and miss is shown below in full. Nothing cherry-picked.

REAL-TIME · NON-STOP
models: 4 × qwen2.5-7B (Q4) context: 8k each recall: 0verload hardware: one consumer GPU per bot
Give your agents this memory → start free trial
ANSWERS LANDING NOW — ALL LANES 0 answered · updating live
connecting to the live stream…
loading benchmark stream…