arXiv:2603.29318cs.AI2026-03被引 4

首个专为手机界面代理个性化能力设计的评测基准。

PSPA-Bench: A Personalized Benchmark for Smartphone GUI Agent

  • 构建12855条真实用户行为指令,覆盖22个应用和10类日常场景。
  • 现有最强代理在个性化任务中成功率有限,表现远未达标。
  • 强调推理模型、感知能力和长期记忆对个性化的关键作用。

智能手机图形用户界面(GUI)代理通过直接操作应用界面执行任务,为实现广泛功能提供无需深度系统集成的路径。然而,现实中的手机使用高度个性化:用户采用多样化的操作流程与偏好,要求代理提供定制化协助而非通用解决方案。现有GUI代理评测基准难以充分捕捉个性化维度,原因在于用户特定数据稀疏及缺乏细粒度评估指标。为此,我们提出PSPA-Bench,首个专注于评估手机GUI代理个性化能力的基准。PSPA-Bench包含超过12,855条与真实用户行为一致的个性化指令,覆盖10个代表性日常使用场景和22个主流移动应用,并引入结构感知的过程评估方法,实现对代理个性化能力的细粒度测量。基于此,我们对11个顶尖GUI代理进行了评测。结果表明,当前方法在个性化设置下表现不佳,即使最强代理也仅取得有限成功。分析进一步揭示三个发展方向:(1) 以推理为导向的模型持续优于通用大语言模型;(2) 感知能力虽简单但至关重要;(3) 反思与长期记忆机制是提升适应性的关键。这些发现确立了PSPA-Bench作为个性化GUI代理系统研究与未来进展的基础。

原文摘要 · Abstract (English)

Smartphone GUI agents execute tasks by operating directly on app interfaces, offering a path to broad capability without deep system integration. However, real-world smartphone use is highly personalized: users adopt diverse workflows and preferences, challenging agents to deliver customized assistance rather than generic solutions. Existing GUI agent benchmarks cannot adequately capture this personalization dimension due to sparse user-specific data and the lack of fine-grained evaluation metrics. To address this gap, we present PSPA-Bench, the benchmark dedicated to evaluating personalization in smartphone GUI agents. PSPA-Bench comprises over 12,855 personalized instructions aligned with real-world user behaviors across 10 representative daily-use scenarios and 22 mobile apps, and introduces a structure-aware process evaluation method that measures agents' personalized capabilities at a fine-grained level. Through PSPA-Bench, we benchmark 11 state-of-the-art GUI agents. Results reveal that current methods perform poorly under personalized settings, with even the strongest agent achieving limited success. Our analysis further highlights three directions for advancing personalized GUI agents: (1) reasoning-oriented models consistently outperform general LLMs, (2) perception remains a simple yet critical capability, and (3) reflection and long-term memory mechanisms are key to improving adaptation. Together, these findings establish PSPA-Bench as a foundation for systematic study and future progress in personalized GUI agents.

GUI代理个性化评测基准手机智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。