arXiv:2410.13757cs.MAcs.AI2024-10NAACL被引 14

让手机助手更懂界面变化,自动纠错并高效完成复杂操作

MobA: Multifaceted Memory-Enhanced Adaptive Planning for Efficient Mobile Task Automation

论文配图:MobA: Multifaceted Memory-Enhanced Adaptive Planning for Efficient Mobile Task Automation
图 1 · 摘自论文原文
  • 引入自适应规划与多面记忆,动态调整任务策略
  • 在真实数据集上任务成功率超基线模型18.6%以上
  • 适合开发智能自动化工具或研究移动AI代理的开发者

基于多模态大语言模型(MLLM)的移动助手在处理设备上的复杂图形用户界面(GUI)交互时面临挑战,这源于GUI环境的动态性与结构复杂性,包括文本、图像及空间关系的融合,以及不同页面和任务间动作空间的差异。为此,我们提出MobA,一种新型的MLLM驱动的移动端助手系统。MobA包含一个具备反思机制的自适应规划模块,可实现错误恢复,并根据实际环境上下文和动作模块执行能力动态调整计划;同时,一个多面记忆模块提供全面的记忆支持,提升系统的适应性与效率。我们还构建了MobBench数据集,用于评估复杂移动端交互。在MobBench和AndroidArena上的实验表明,MobA能够有效应对动态GUI环境,完成复杂移动端任务。

原文摘要 · Abstract (English)

Existing Multimodal Large Language Model (MLLM)-based agents face significant challenges in handling complex GUI (Graphical User Interface) interactions on devices. These challenges arise from the dynamic and structured nature of GUI environments, which integrate text, images, and spatial relationships, as well as the variability in action spaces across different pages and tasks. To address these limitations, we propose MobA, a novel MLLM-based mobile assistant system. MobA introduces an adaptive planning module that incorporates a reflection mechanism for error recovery and dynamically adjusts plans to align with the real environment contexts and action module's execution capacity. Additionally, a multifaceted memory module provides comprehensive memory support to enhance adaptability and efficiency. We also present MobBench, a dataset designed for complex mobile interactions. Experimental results on MobBench and AndroidArena demonstrate MobA's ability to handle dynamic GUI environments and perform complex mobile tasks.

移动自动化大模型自适应规划多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。