构建可主动模拟用户行为的环境,评估智能助手的预判能力。
Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants

- 将应用建模为带状态的有限自动机,支持真实交互序列模拟。
- 设计包含143项任务的基准测试,覆盖沟通、办公等多场景。
- 适合研究主动式助手的算法与评估方法的团队使用。
能够预判用户需求并自主执行任务的主动式智能助手具有巨大潜力,但缺乏真实的用户模拟框架限制了其发展。现有方法将应用程序建模为无状态的工具调用API,无法捕捉数字环境中用户交互的状态性与顺序性,导致真实用户模拟不可行。我们提出主动代理研究环境(Pare),一个用于在数字环境中构建和评估主动代理的框架。Pare将应用建模为带有状态转移和状态相关动作空间的有限状态机,实现主动用户模拟。在此基础上,我们构建了Pare-Bench基准,包含143个涵盖通信、生产力、日程安排和生活方式类应用的多样化任务,旨在测试上下文感知、目标推断、干预时机把握及多应用协同能力。
原文摘要 · Abstract (English)
Proactive agents that anticipate user needs and autonomously execute tasks hold great promise as digital assistants, yet the lack of realistic user simulation frameworks hinders their development. Existing approaches model apps as flat tool-calling APIs, failing to capture the stateful and sequential nature of user interaction in digital environments and making realistic user simulation infeasible. We introduce Proactive Agent Research Environment (Pare), a framework for building and evaluating proactive agents in digital environments. Pare models applications as finite state machines with stateful navigation and state-dependent action space for the user simulator, enabling active user simulation. Building on this foundation, we present Pare-Bench, a benchmark of 143 diverse tasks spanning communication, productivity, scheduling, and lifestyle apps, designed to test context observation, goal inference, intervention timing, and multi-app orchestration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。