构建更真实全面的移动端视觉语言模型评估基准,支持多路径、噪声环境与主动交互测试。
Mobile-Bench-v2: A More Realistic and Comprehensive Benchmark for VLM-based Mobile Agents
- 采用基于槽位的指令生成法,构建多场景评估数据集
- 包含多路径、噪声弹窗、模糊指令等12个子任务,覆盖真实使用场景
- 适合评估智能助手在复杂手机环境中的鲁棒性与主动性
基于视觉语言模型的移动代理因其能与智能手机GUI和结构化文本交互而日益流行,可完成日常任务。然而现有在线基准因环境动态变化难以获得稳定奖励信号;离线基准仅通过单路径轨迹评估,与图形界面任务固有的多解特性不符。此外,两类基准均缺乏对噪声处理能力或主动交互能力的评估,因未引入干扰应用或过于完整的指令。为此,我们提出基于槽位的指令生成方法,构建更真实、全面的Mobile-Bench-v2基准。该基准包含通用任务划分,采用离线多路径评估以检验任务执行中步骤奖励获取能力;设有基于弹窗与广告应用的噪声子集,以及名为AITZ-Noise的污染环境子集;还发布含预设问答交互的模糊指令子集,用于评估代理的主动交互能力。我们在AppAgent-v1单代理框架、Mobile-Agent-v2多代理框架及其他移动代理如UI-Tars、OS-Atlas上进行评测。代码与数据已公开于https://huggingface.co/datasets/xwk123/MobileBench-v2。
原文摘要 · Abstract (English)
VLM-based mobile agents are increasingly popular due to their capabilities to interact with smartphone GUIs and XML-structured texts and to complete daily tasks. However, existing online benchmarks struggle with obtaining stable reward signals due to dynamic environmental changes. Offline benchmarks evaluate the agents through single-path trajectories, which stands in contrast to the inherently multi-solution characteristics of GUI tasks. Additionally, both types of benchmarks fail to assess whether mobile agents can handle noise or engage in proactive interactions due to a lack of noisy apps or overly full instructions during the evaluation process. To address these limitations, we use a slot-based instruction generation method to construct a more realistic and comprehensive benchmark named Mobile-Bench-v2. Mobile-Bench-v2 includes a common task split, with offline multi-path evaluation to assess the agent's ability to obtain step rewards during task execution. It contains a noisy split based on pop-ups and ads apps, and a contaminated split named AITZ-Noise to formulate a real noisy environment. Furthermore, an ambiguous instruction split with preset Q\&A interactions is released to evaluate the agent's proactive interaction capabilities. We conduct evaluations on these splits using the single-agent framework AppAgent-v1, the multi-agent framework Mobile-Agent-v2, as well as other mobile agents such as UI-Tars and OS-Atlas. Code and data are available at https://huggingface.co/datasets/xwk123/MobileBench-v2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。