arXiv:2501.01149cs.AI2025-01ACL被引 9

构建动态安卓应用评估体系,让手机AI代理真实任务表现可测。

A3: Android Agent Arena for Mobile GUI Agents with Essential-State Procedural Evaluation

  • 以关键状态为依据,用多模态大模型做奖励模型逐步验证任务完成。
  • 覆盖20类100个真实在线应用任务,涵盖20个热门谷歌商店应用。
  • 提供设备交互、环境重置工具,支持人与智能体数据采集。

大型语言模型(LLMs)和多模态大语言模型(MLLMs)的发展推动了移动端图形用户界面(GUI)AI代理的兴起,使其能够自主执行移动设备上的任务。然而,当前移动GUI代理评估存在显著空白:现有基准大多依赖静态画面评估(如AndroidControl)或离线静态应用(如AndroidWorld),难以反映代理在动态真实在线应用中的表现。为此,我们提出Android Agent Arena(A3),一个基于“关键状态”的程序化评估系统。A3构建了一个由20个广泛使用的动态在线应用组成的基准,涵盖20个类别,包含100个任务,确保评估全面性。其创新的“关键状态”评估方法利用MLLM作为奖励模型,逐步验证任务完成与过程达成,克服了传统功能评估在动态在线应用中的局限。此外,A3还提供工具包,用于简化安卓设备交互、重置在线环境与应用,并支持从人类与代理演示中收集数据。A3完整系统(含基准与工具)将公开发布,为移动端GUI代理的研究与发展提供坚实基础。

原文摘要 · Abstract (English)

The advancement of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) has catalyzed the development of mobile graphic user interface (GUI) AI agents, which is designed to autonomously perform tasks on mobile devices. However, a significant gap persists in mobile GUI agent evaluation, where existing benchmarks predominantly rely on either static frame assessments such as AndroidControl or offline static apps such as AndroidWorld and thus fail to capture agent performance in dynamic, real-world online mobile apps. To address this gap, we present Android Agent Arena (A3), a novel "essential-state" based procedural evaluation system for mobile GUI agents. A3 introduces a benchmark of 100 tasks derived from 20 widely-used, dynamic online apps across 20 categories from the Google Play Store, ensuring evaluation comprehension. A3 also presents a novel "essential-state" based procedural evaluation method that leverages MLLMs as reward models to progressively verify task completion and process achievement. This evaluation approach address the limitations of traditional function based evaluation methods on online dynamic apps. Furthermore, A3 includes a toolkit to streamline Android device interaction, reset online environment and apps and facilitate data collection from both human and agent demonstrations. The complete A3 system, including the benchmark and tools, will be publicly released to provide a robust foundation for future research and development in mobile GUI agents.

移动AIGUI代理评估基准多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。