测试大模型在模拟一周中延迟执行任务的能力,发现最佳表现仅65.1%。
PM-Bench: Evaluating Prospective Memory in LLM Agents

- 设计七日模拟环境,评估模型维持并适时执行延迟指令的能力。
- 最先进模型在基准上最高仅达65.1%的F1分数,普遍表现不佳。
- 适合研究大模型记忆与任务调度机制的开发者和研究人员。
当前智能体人工智能面临的核心挑战之一是前瞻性记忆:即在持续进行其他活动时,于特定未来触发条件或状态下执行既定意图。我们提出PM-Bench,一个基于文本的基准测试,用于衡量现代大语言模型智能体的前瞻性记忆能力。受认知科学中虚拟一周范式启发,该基准评估模型维护用户意图、执行延迟任务及监测潜在环境变化的能力。在模拟七天周期中,智能体需持续进行当前任务,同时判断是否有待办事项到期。我们在八种不同智能体配置下对比了八种前沿大模型的表现。结果表明,所有设置下任务均具挑战性:最佳方法——GPT-5.4智能体——在评估中仅达到65.1%的F1分数。此外,无单一策略能在所有模型中统一提升表现。我们公开发布PM-Bench,作为诊断失败原因并开发训练或推理阶段干预措施的可控测试平台,以支持可靠前瞻行为。
原文摘要 · Abstract (English)
A significant challenge in agentic AI is prospective memory: the ability to execute an intention at a specific future cue or state while other activities are ongoing. We introduce PM-Bench, a text-based benchmark for measuring prospective memory capabilities in modern LLM agents. Inspired by the Virtual Week paradigm from cognitive science, PM-Bench evaluates how well LLM agents maintain user intentions, execute delayed intentions, and monitor latent environment changes. Over the course of a simulated seven-day week, agents must continue an ongoing activity while deciding whether any deferred task is due. We compare eight state-of-the-art LLMs on PM-Bench under eight different agent configurations. PM-Bench proves challenging across all settings: the best method, a GPT-5.4 agent, reaches only 65.1\% F1 score under our evaluation. Furthermore, no single strategy for improving prospective memory dominates across models. We release PM-Bench as a controlled testbed for diagnosing these failures and developing training or inference-time interventions that support reliable prospective behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。