首个面向主动智能体的事件感知评测基准,测试其预判用户日程的能力。
ProEvent: An Event-centric Benchmark for Proactive Agents

- 基于聊天记录动态推断用户待办事件,实现事件感知与追踪
- 现有大模型在事件响应中正确率仅26.7%,且频繁误动作
- 适合研究主动智能体、具身认知与人机协作的学者参考
主动智能体需在无明确指令下感知环境上下文,提前预判用户需求并提供自主协助。其核心能力之一是识别与追踪用户的未来事件,从而实现持续、场景化的支持。例如,通过记录计划徒步的时间与地点,智能体可提前推送天气提醒或出发前导航支持。然而,现有研究普遍忽视以事件为中心的主动服务,且主动行为的开放性给可靠评估带来挑战。为此,我们提出ProEvent——首个面向事件感知的主动智能体评测基准,旨在评估智能体基于实时聊天对话维护用户日程的能力。ProEvent包含合成但真实感强的对话数据,涵盖用户间动态交互、并发会话及现实噪声。评测维度包括响应时机、单步正确性与多步正确性。在8个LLM与流程上的实验表明,当前智能体常出现过度反应,且难以处理事件取消。值得注意的是,即使GPT-5.1在26.7%的场景中仍能正确响应。定性分析进一步揭示当前大模型作为主动代理的根本局限,尤其在识别隐含事件和从用户第一人称视角推理方面。
原文摘要 · Abstract (English)
Proactive agents are expected to anticipate user needs and provide autonomous assistance by perceiving environmental context without explicit instructions. A fundamental capability of such agents is to identify and track users' upcoming events, enabling continuous and event-specific assistance. For example, by recording the time and location of a planned hike, an agent can deliver weather reminders in advance or provide navigation support before departure. However, existing works on proactive agents largely overlook event-centric assistance, and the open-ended nature of proactive assistance poses challenges for reliable evaluation. To bridge these gaps, we introduce ProEvent, the first event-centric benchmark designed to assess an agent's ability to proactively maintain a user's timetable based on ongoing instant messaging chats. ProEvent provides synthesized yet realistic chats that consider the dynamic interaction among users, concurrent chat threads, and noise in the real world, and evaluates proactive agents on response timing, single-step response correctness, and multi-step response correctness. Experiments on eight LLMs and pipelines reveal that current agents frequently overact and struggle with event cancellation. Notably, even GPT-5.1 only reacts correctly in 26.7% of scenarios. Further qualitative analysis reveals fundamental limitations of current LLMs as proactive agents, particularly in detecting implicit events and reasoning from the user's first-person perspective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。