测试智能体跨会话记忆与推理能力,发现现有模型常误用历史信息。
Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations

- 设计多会话任务评估框架,要求智能体持续记忆用户状态。
- 当前智能体因误判历史为最新状态,导致任务失败率超70%。
- 适合研究长期交互、个性化服务的AI开发者参考。
近期代理型AI已能通过工具使用、推理和多步规划完成复杂任务。然而,现有基准测试仅限单次会话,忽略了代理需整合过往行为、用户偏好和先前决策以实现个性化目标这一关键因素。我们提出Momento,一个用于多会话服务环境中持久性代理任务完成的基准,要求代理在执行有影响的工具化操作时,处理时间依赖性和随会话演进的用户目标。实验结果表明,当前代理主要因错误估计用户状态而失败:将前一会话历史视为当前上下文的可靠代理,而非需重新验证的过时信息,暴露出当前代理能力与真实长周期人机交互需求之间存在显著差距。
原文摘要 · Abstract (English)
Recent advances in agentic AI have enabled agents to complete complex tasks through tool use, reasoning, and multi-step planning. Yet existing benchmarks evaluate agents within a single session, ignoring past actions, stated preferences, and prior decisions that agents must integrate to fulfill personalized user goals. We introduce Momento, a benchmark for persistent agentic task completion in multi-session service environments, requiring agents to take consequential, tool-mediated actions while resolving temporal dependencies and evolving user goals across sessions. Experimental results reveal that current agents fail primarily through misestimation of user state, treating prior session history as a reliable proxy for current context rather than stale information requiring re-validation, highlighting a substantial gap between current agent capabilities and realistic long-horizon human-agent interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。