构建真实用户场景下的个人助手评估基准,测试其跨服务状态管理能力。
PAUSE: A User-Centric Benchmark for Personal AI Assistants in Unified Service Environments

- 设计多轮交互的用户模拟系统,评测助手在状态与权限约束下的表现。
- 顶尖模型在需状态推理的任务上仅完成70%以下,暴露普遍性失败模式。
- 适用于研究个人助理长期交互、配置感知与多服务协调的学者与开发者。
个人AI助手正作为面向任务、工具增强的智能体,在统一服务环境中支持日常活动。在真实场景中,助手需持续维护用户状态、遵守个性化配置与权限,并在多服务间进行长周期、约束感知的交互。现有基准常割裂服务上下文或忽略用户状态,难以评估真实环境下的用户中心行为。本文提出PAUSE,一个面向有状态、服务集成环境的用户中心评估基准。该基准通过要求智能体在异构用户资源间协调动作,同时保持与环境状态和授权约束的一致性,捕捉现实部署的核心挑战。引入真实用户模拟实现用户-代理交互,超越静态工具执行评估。采用多范式评估框架:开放性服务管理任务使用语义与轨迹级行为指标,约束密集型任务则通过基于状态的确定性验证。结果表明,即使最先进的专有模型在需状态推理与配置感知的任务上也未达70%任务完成率,揭示出一致且可解释的失败模式。最后,提出一种用户中心合成流水线,支持可扩展生成一致的服务环境、用户配置与可靠标注任务,促进基准演进与后续研究。
原文摘要 · Abstract (English)
Personal AI assistants are increasingly deployed as task-oriented, tool-augmented agents that operate within unified service environments to support everyday user activities. In realistic settings, such assistants must reason over persistent user state, respect user-specific configurations and permissions, and sustain long-horizon, constraint-aware interactions across multiple services. Existing benchmarks, however, often fragment service contexts or abstract away user state, limiting their ability to evaluate user-centric personal assistant behavior in realistic service settings. We introduce PAUSE, a user-centric benchmark for evaluating personal AI assistants in stateful, service-integrated environments. PAUSE captures core challenges of real-world assistant deployment by requiring agents to coordinate actions across heterogeneous user-owned resources while maintaining consistency with environment state, authorization constraints over multi-turn interactions. The benchmark incorporates explicit user-agent interaction via realistic user simulation, enabling evaluation beyond static tool execution. To support principled and reproducible evaluation, PAUSE adopts a multi-regime evaluation framework aligned with task characteristics. Open-ended service management tasks are assessed using semantic and trajectory-level behavioral metrics, while constraint-intensive tasks admit deterministic, state-based verification. Benchmark results show that even state-of-the-art proprietary models fail to reach 70% task completion on scenarios requiring stateful reasoning and configuration awareness, revealing consistent and interpretable failure patterns. Finally, we present a user-centric synthesis pipeline that enables scalable generation of coherent service environments, user configurations, and reliably annotated tasks, supporting benchmark extensibility and future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。