[ICLR 2026] Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
-
Updated
Aug 31, 2026 - Python
[ICLR 2026] Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
First-Person Agent Memory Bench. 10 Categories including fact recall, multi-hop links, temporal reasoning, fact overwrites, speaker traps, refusal, credibility, and agentic tool usage. 540K token / 60 session corpus, all in first person. Dynamic output-answer-key portion. Comprehensive report with visuals and miss breakdown.
Reduce web pages to clean text before they reach your AI agent’s context window, saving tokens and cutting noise
To associate your repository with the beam-benchmark topic, visit your repo's landing page and select "manage topics."