arXiv:2509.02444cs.AIcs.CL2025-09被引 3

打造通用高效手机智能体,能跨应用跨设备完成复杂任务。

AppCopilot: Toward General, Accurate, Long-Horizon, and Efficient Mobile Agent

  • 多模态多智能体架构,融合思维链与分层规划提升决策能力。
  • 在跨应用长序列任务中准确率超90%,运行效率提升3倍以上。
  • 适合开发智能助手、自动化工具的工程师和研究者参考。

随着大语言模型和多模态模型的快速发展,移动智能体领域迅速扩张却未解决核心挑战。本文指出四大关键问题:任务、应用与设备间的泛化能力;精确的屏幕交互与点击定位;长期多步目标的持续执行能力;以及资源受限设备上的高效运行。我们提出AppCopilot,一个端到端的通用型移动智能体系统,覆盖数据采集、训练、微调、高效推理及移动端部署全流程。模型层整合支持中英文的多模态基础模型;推理与控制层结合思维链、分层任务规划与多智能体协作;执行层实现经验自适应、语音交互、函数调用、跨应用/跨设备编排及全面的手机应用支持。系统通过性能分析驱动优化,在异构硬件上实现低延迟与低内存占用。实验表明,AppCopilot在泛化性、屏幕操作精度、长周期任务完成率和运行效率上均有显著提升。本文提出统一框架与可落地的技术路线,为通用移动智能体提供明确发展蓝图。

原文摘要 · Abstract (English)

With the raid evolution of large language models and multimodal models, the mobile-agent landscape has proliferated without converging on the fundamental challenges. This paper identifies four core problems that should be solved for mobile agents to deliver practical, scalable impact: (1) generalization across tasks, APPs, and devices; (2) accuracy, specifically precise on-screen interaction and click targeting; (3) long-horizon capability for sustained, multi-step goals; and (4) efficiency, specifically high-performance runtime on resource-constrained devices. We present AppCopilot, a multimodal, multi-agent, general-purpose mobile agent that operates across applications. AppCopilot operationalizes this position through an end-to-end pipeline spanning data collection, training, finetuning, efficient inference, and PC/mobile application. At the model layer, it integrates multimodal foundation models with robust Chinese-English support. At the reasoning and control layer, it combines chain-of-thought reasoning, hierarchical task planning and decomposition, and multi-agent collaboration. At the execution layer, it enables experiential adaptation, voice interaction, function calling, cross-APP and cross-device orchestration, and comprehensive mobile APP support. The system design incorporates profiling-driven optimization for latency and memory across heterogeneous hardware. Empirically, AppCopilot achieves significant improvements on four dimensions: stronger generalization, higher precision of on screen actions, more reliable long horizon task completion, and faster, more resource efficient runtime. By articulating a cohesive position and a reference architecture that closes the loop from data collection, training to finetuning and efficient inference, this paper offers a concrete roadmap for general purpose mobile agent and provides actionable guidance.

移动智能体多智能体长程任务效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。