First-Person Agent Memory Bench. 10 Categories including fact recall, multi-hop links, temporal reasoning, fact overwrites, speaker traps, refusal, credibility, and agentic tool usage. 540K token / 60 session corpus, all in first person. Dynamic output-answer-key portion. Comprehensive report with visuals and miss breakdown.
-
Updated
Aug 27, 2026 - Python