构建大规模多模态数据集,提升手机控制智能体训练与评估效果
DigiData: Training and Evaluating General-Purpose Mobile Control Agents
- 基于应用功能深度探索构建多样化目标数据集
- 提出动态评估协议和AI辅助评测替代传统步数准确率
- 适合研究人机交互、智能体训练的开发者和研究人员
能够控制用户界面的AI智能体有望重塑人机交互方式。为加速这一进程,高质量数据集和可靠评估方法至关重要。本文提出DigiData,一个大规模、高质、多样、多模态的移动端控制智能体训练数据集。与现有数据集依赖非结构化交互不同,DigiData通过全面探索应用功能构建目标,显著提升目标多样性与复杂性。同时,我们推出DigiData-Bench基准,用于评估智能体在真实复杂任务上的表现。实验表明,常用步数准确率无法可靠评估移动端智能体性能,因此我们提出动态评估协议和AI驱动评估作为更严格的替代方案。本工作旨在推动移动端控制智能体的发展,实现更自然高效的人机交互。
原文摘要 · Abstract (English)
AI agents capable of controlling user interfaces have the potential to transform human interaction with digital devices. To accelerate this transformation, two fundamental building blocks are essential: high-quality datasets that enable agents to achieve complex and human-relevant goals, and robust evaluation methods that allow researchers and practitioners to rapidly enhance agent performance. In this paper, we introduce DigiData, a large-scale, high-quality, diverse, multi-modal dataset designed for training mobile control agents. Unlike existing datasets, which derive goals from unstructured interactions, DigiData is meticulously constructed through comprehensive exploration of app features, resulting in greater diversity and higher goal complexity. Additionally, we present DigiData-Bench, a benchmark for evaluating mobile control agents on real-world complex tasks. We demonstrate that the commonly used step-accuracy metric falls short in reliably assessing mobile control agents and, to address this, we propose dynamic evaluation protocols and AI-powered evaluations as rigorous alternatives for agent assessment. Our contributions aim to significantly advance the development of mobile control agents, paving the way for more intuitive and effective human-device interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。