Nobody would point a model with an 8,000-token window at a benchmark whose conversations run to a million tokens and beyond. Who would even try that? We run four of them at once — tiny qwen-7B bots, one per consumer GPU — and they pass anyway. Because the model isn’t doing the remembering: 0verload is. The model reasons over what it’s handed; the store hands it back the right fact in milliseconds. The number that matters is recall latency — and every question, answer, and miss is shown below in full. Nothing cherry-picked.