arXiv:2503.05143cs.AIcs.LG2025-03被引 3

首个面向异构用户数据的移动智能体联邦训练基准平台。

FedMABench: Benchmarking Mobile Agents on Decentralized Heterogeneous User Data

  • 构建首个专为异构场景设计的移动智能体联邦学习评测基准。
  • 涵盖800+应用、30多个数据子集,验证联邦算法优于本地训练。
  • 适合研究移动智能体、联邦学习与跨设备协同的开发者使用。

移动智能体近年来受到广泛关注。传统移动智能体训练依赖集中式数据收集,成本高且难以扩展。利用联邦学习分布训练可借助真实用户数据,提升可扩展性并降低成本。然而,缺乏标准化基准严重制约该领域进展。为此,我们提出FedMABench,首个针对异构场景的移动智能体联邦训练与评估基准。该基准包含6个数据集、30多个子集、8种联邦算法、10余种基础模型,覆盖5大类超过800个应用,提供全面的评测框架。大量实验揭示:联邦算法在多数情况下优于本地训练;特定应用的数据分布对异构性影响显著;甚至不同类别应用在训练中也表现出相关性。代码与数据已公开于https://github.com/wwh0411/FedMABench及https://huggingface.co/datasets/wwh0411/FedMABench。

原文摘要 · Abstract (English)

Mobile agents have attracted tremendous research participation recently. Traditional approaches to mobile agent training rely on centralized data collection, leading to high cost and limited scalability. Distributed training utilizing federated learning offers an alternative by harnessing real-world user data, providing scalability and reducing costs. However, pivotal challenges, including the absence of standardized benchmarks, hinder progress in this field. To tackle the challenges, we introduce FedMABench, the first benchmark for federated training and evaluation of mobile agents, specifically designed for heterogeneous scenarios. FedMABench features 6 datasets with 30+ subsets, 8 federated algorithms, 10+ base models, and over 800 apps across 5 categories, providing a comprehensive framework for evaluating mobile agents across diverse environments. Through extensive experiments, we uncover several key insights: federated algorithms consistently outperform local training; the distribution of specific apps plays a crucial role in heterogeneity; and, even apps from distinct categories can exhibit correlations during training. FedMABench is publicly available at: https://github.com/wwh0411/FedMABench with the datasets at: https://huggingface.co/datasets/wwh0411/FedMABench.

联邦学习移动智能体基准测试异构数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。