arXiv:2410.24024cs.AI2024-10被引 104

构建可复现的安卓智能体评测框架,提升自动化操作成功率。

AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents

  • 统一环境支持语言与多模态模型在相同操作空间中训练
  • 138项任务覆盖9个应用,推动平均成功率提升至21.50%(LLM)和13.28%(LMM)
  • 开源框架支持研究者系统评估安卓智能体,适合智能交互方向从业者

自主智能体在现实世界交互中日益重要,尤其是安卓平台上的智能体。然而,现有针对安卓智能体的训练与评估研究缺乏对开源与闭源模型的系统性探索。本文提出AndroidLab,一个系统的安卓智能体框架,包含支持多模态、具统一动作空间的操作环境及可复现的基准测试。该基准涵盖138项任务,分布在9个基于虚拟设备的应用中。利用AndroidLab环境,我们构建了安卓指令数据集,并训练了6个开源大语言模型(LLMs)与多模态模型(LMMs),使LLM平均成功率达21.50%(原4.59%),LMM达13.28%(原1.93%)。AndroidLab已开源,可通过https://github.com/THUDM/Android-Lab获取。

原文摘要 · Abstract (English)

Autonomous agents have become increasingly important for interacting with the real world. Android agents, in particular, have been recently a frequently-mentioned interaction method. However, existing studies for training and evaluating Android agents lack systematic research on both open-source and closed-source models. In this work, we propose AndroidLab as a systematic Android agent framework. It includes an operation environment with different modalities, action space, and a reproducible benchmark. It supports both large language models (LLMs) and multimodal models (LMMs) in the same action space. AndroidLab benchmark includes predefined Android virtual devices and 138 tasks across nine apps built on these devices. By using the AndroidLab environment, we develop an Android Instruction dataset and train six open-source LLMs and LMMs, lifting the average success rates from 4.59% to 21.50% for LLMs and from 1.93% to 13.28% for LMMs. AndroidLab is open-sourced and publicly available at https://github.com/THUDM/Android-Lab.

智能体安卓评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。