MemoryArena: Agent Memory Evaluation Shifts from Text Recall to Cross-Session Action
Submitted on February 18, 2026, MemoryArena constructs a cross-session, subtask-interdependent Memory-Agent-Environment evaluation, showing that agents near saturation on existing long-context memory benchmarks still perform poorly when required to use past actions and feedback to guide subsequent tasks.
Traditional memory benchmarks often measure 'remembering information' and 'completing actions' separately, easily overestimating the long-term capability of real agents. MemoryArena requires the system to form memories in prior interactions and then apply those experiences to subsequent planning, search, and reasoning.
The benchmark covers web navigation, preference-constrained planning, progressive information search, and sequential formal reasoning, and designs interdependent multi-session subtasks. The evaluation simultaneously examines recall accuracy, memory distillation, cross-session transfer, continuous planning, and final action outcomes.
Agents for customer service, research, personal assistants, and enterprise processes cannot rely solely on static QA memory scores to prove long-term reliability; acceptance also needs to cover state changes, error feedback, and cross-session task outcomes.
When testing agents, enterprises should add cross-day tasks and regression sets that depend on prior feedback, replacing simple historical information recall rates with whether the final business action is correct.
Next steps should focus on task coverage, environment leakage, review reproducibility, and the trade-offs of different memory architectures in latency, cost, privacy, and long-term error propagation.