首个评估移动界面代理在模糊指令下双向意图对齐能力的基准
AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
- 构建四类指令清晰度分类体系,模拟真实用户表达模糊场景
- 240个生态有效任务测试表明,主动交互可显著提升任务成功率
- 开发自动化评价框架MUSE,与人工判断高度一致
评估移动图形界面代理进展离不开基准测试。现实中用户常无法一次性完整表述任务需求,指令往往模糊不清。因此,代理需通过主动追问与交互逐步理解真实意图。但现有基准多假设指令完整明确,仅评估单轮执行,忽略代理的意图对齐能力。为此,我们提出AmbiBench,首个引入指令清晰度分类的基准,将评估从单向指令遵循转向双向意图对齐。基于认知差距理论,我们设计四种清晰度等级:详细、标准、不完整、模糊。构建包含25个应用、240个生态有效任务的严格数据集。为支持动态环境评估,开发MUSE(移动用户满意度评估器),一个基于多智能体大模型判官的自动化框架,从结果有效性、执行质量、交互质量三维度进行细粒度审计。实证结果显示,当前最优代理在不同清晰度下的性能边界,量化了主动交互带来的增益,并验证了MUSE与人类判断的高度相关性。本工作重新定义了评估标准,为下一代真正理解用户意图的智能体奠定基础。
原文摘要 · Abstract (English)
Benchmarks are paramount for gauging progress in the domain of Mobile GUI Agents. In practical scenarios, users frequently fail to articulate precise directives containing full task details at the onset, and their expressions are typically ambiguous. Consequently, agents are required to converge on the user's true intent via active clarification and interaction during execution. However, existing benchmarks predominantly operate under the idealized assumption that user-issued instructions are complete and unequivocal. This paradigm focuses exclusively on assessing single-turn execution while overlooking the alignment capability of the agent. To address this limitation, we introduce AmbiBench, the first benchmark incorporating a taxonomy of instruction clarity to shift evaluation from unidirectional instruction following to bidirectional intent alignment. Grounded in Cognitive Gap theory, we propose a taxonomy of four clarity levels: Detailed, Standard, Incomplete, and Ambiguous. We construct a rigorous dataset of 240 ecologically valid tasks across 25 applications, subject to strict review protocols. Furthermore, targeting evaluation in dynamic environments, we develop MUSE (Mobile User Satisfaction Evaluator), an automated framework utilizing an MLLM-as-a-judge multi-agent architecture. MUSE performs fine-grained auditing across three dimensions: Outcome Effectiveness, Execution Quality, and Interaction Quality. Empirical results on AmbiBench reveal the performance boundaries of SoTA agents across different clarity levels, quantify the gains derived from active interaction, and validate the strong correlation between MUSE and human judgment. This work redefines evaluation standards, laying the foundation for next-generation agents capable of truly understanding user intent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。