构建交互式移动助理评估基准,测试其个性化与主动服务能力。
KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation
- 通过动态用户模拟器实现多轮偏好获取与主动干预决策测试。
- 前沿模型在模糊指令下表现下降至50%以下,主因是偏好获取不足。
- 适合研究个性化智能助手、人机交互与主动服务的开发者与学者。
个性化移动代理需推断用户偏好并主动提供适配帮助,但现有基准无法捕捉这一需求。以往工作仅评估静态历史中的偏好恢复或固定上下文下的意图预测,未检验代理是否能通过交互获取缺失偏好,或在实时图形界面中判断何时介入、征求同意或保持沉默。本文提出KnowU-Bench,一个基于可复现Android模拟环境的在线评估基准,涵盖42项通用GUI任务、86项个性化任务和64项主动任务。不同于将用户偏好视为静态上下文的做法,KnowU-Bench隐藏用户档案,仅暴露行为日志,迫使代理进行真实偏好推断而非简单查表。为支持多轮偏好澄清,引入基于结构化档案的LLM驱动用户模拟器,实现自然对话与主动同意处理。此外,基准覆盖从界面执行、同意协商到拒绝后克制的完整主动决策链,采用规则验证与LLM-as-a-Judge混合协议评估。实验表明,即使前沿模型如Claude Sonnet 4.6,在需要偏好推断或干预校准的模糊指令下,性能也低于50%,核心瓶颈并非界面操作,而是偏好获取与干预校准能力,揭示了高效界面操作与可信个性化服务之间的根本差距。
原文摘要 · Abstract (English)
Personalized mobile agents that infer user preferences and calibrate proactive assistance hold great promise as everyday digital assistants, yet existing benchmarks fail to capture what this requires. Prior work evaluates preference recovery from static histories or intent prediction from fixed contexts. Neither tests whether an agent can elicit missing preferences through interaction, nor whether it can decide when to intervene, seek consent, or remain silent in a live GUI environment. We introduce KnowU-Bench, an online benchmark for personalized mobile agents built on a reproducible Android emulation environment, covering 42 general GUI tasks, 86 personalized tasks, and 64 proactive tasks. Unlike prior work that treats user preferences as static context, KnowU-Bench hides the user profile from the agent and exposes only behavioral logs, forcing genuine preference inference rather than context lookup. To support multi-turn preference elicitation, it instantiates an LLM-driven user simulator grounded in structured profiles, enabling realistic clarification dialogues and proactive consent handling. Beyond personalization, KnowU-Bench provides comprehensive evaluation of the complete proactive decision chain, including grounded GUI execution, consent negotiation, and post-rejection restraint, evaluated through a hybrid protocol combining rule-based verification with LLM-as-a-Judge scoring. Our experiments reveal a striking degradation: agents that excel at explicit task execution fall below 50% under vague instructions requiring user preference inference or intervention calibration, even for frontier models like Claude Sonnet 4.6. The core bottlenecks are not GUI navigation but preference acquisition and intervention calibration, exposing a fundamental gap between competent interface operation and trustworthy personal assistance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。