arXiv:2502.02982cs.AI2025-02被引 9

用用户自动生成的手机操作数据训练智能助手,成本低且保护隐私

MobileA3gent: Training Mobile GUI Agents Using Decentralized Self-Sourced Data from Diverse Users

  • 自动收集用户日常使用数据,无需人工标注
  • 在仅1%成本下达到传统方法性能,支持跨设备泛化
  • 适合移动端自动化、隐私敏感场景的智能体开发

移动GUI智能体的发展为手机任务自动化带来了新机遇。训练此类智能体需大规模高质量数据,但依赖人工标注成本过高。全球海量手机用户若能实现自动化数据采集,将产生前所未有的数据规模与模型能力。然而面临两大挑战:(1) 无须人工干预地提取用户指令;(2) 在保障隐私前提下利用分布式用户数据。为此,我们提出MobileA3gent,一个基于去中心化自生数据的协作训练框架。该框架包含两部分:(1) Auto-Annotation,可在用户日常使用中低成本自动收集高质量数据集;(2) FedVLM-A,通过结合任务级与步骤级变异性,改进非独立同分布下的联邦视觉语言模型训练。大量实验表明,MobileA3gent以仅1%的成本实现优于传统方法的性能,展现出实际应用潜力。

原文摘要 · Abstract (English)

The advancement of mobile GUI agents has opened new opportunities for automating tasks on mobile devices. Training these agents requires large-scale high-quality data, which is prohibitively expensive when relying on human labor. Given the vast population of global mobile phone users, if automated data collection from them becomes feasible, the resulting data volume and the subsequently trained mobile agents could reach unprecedented levels. Nevertheless, two major challenges arise: (1) extracting user instructions without human intervention and (2) utilizing distributed user data while preserving privacy. To tackle these challenges, we propose MobileA3gent, a collaborative framework that trains mobile GUI Agents using decentralized self-sourced data from diverse users. The framework comprises two components, each targeting a specific challenge: (1) Auto-Annotation, which enables the automatic collection of high-quality datasets during users' routine phone usage with minimal cost. (2) FedVLM-A, which enhances federated VLM training under non-IID distributions by incorporating adapted global aggregation based on both episode-level and step-level variability. Extensive experiments prove that MobileA3gent achieves superior performance over traditional approaches at only 1% of the cost, highlighting its potential for real-world applications

移动智能体联邦学习数据生成隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。