综述多模态移动智能体最新进展,聚焦实时适应与跨模态交互。
Foundations and Recent Trends in Multimodal Mobile Agents: A Survey
- 按提示词与训练两种方式分类当前多模态移动智能体方法
- 新评估基准更精准衡量智能体在动态环境中的表现
- 适合对智能体系统、人机交互感兴趣的科研人员参考
移动智能体在复杂动态环境中执行自动化任务至关重要。随着基础模型的发展,对能实时适应并处理多模态数据的智能体需求日益增长。本文综述移动智能体技术,重点分析提升实时适应性与多模态交互能力的最新进展。近期开发的评估基准更准确反映移动任务的静态与交互环境,可更可靠评估智能体性能。我们将其进展归纳为两类:基于提示的方法(利用大语言模型进行指令驱动的任务执行)和基于训练的方法(微调多模态模型以适配移动场景)。此外,还探讨了增强智能体性能的互补技术。通过分析关键挑战并提出未来方向,本综述为推进移动智能体技术提供重要参考。完整资源列表见 https://github.com/aialt/awesome-mobile-agents
原文摘要 · Abstract (English)
Mobile agents are essential for automating tasks in complex and dynamic mobile environments. As foundation models evolve, the demands for agents that can adapt in real-time and process multimodal data have grown. This survey provides a comprehensive review of mobile agent technologies, focusing on recent advancements that enhance real-time adaptability and multimodal interaction. Recent evaluation benchmarks have been developed better to capture the static and interactive environments of mobile tasks, offering more accurate assessments of agents' performance. We then categorize these advancements into two main approaches: prompt-based methods, which utilize large language models (LLMs) for instruction-based task execution, and training-based methods, which fine-tune multimodal models for mobile-specific applications. Additionally, we explore complementary technologies that augment agent performance. By discussing key challenges and outlining future research directions, this survey offers valuable insights for advancing mobile agent technologies. A comprehensive resource list is available at https://github.com/aialt/awesome-mobile-agents
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。