arXiv:2501.04575cs.AIcs.CL2025-01Conference of the …被引 80

让电脑助手自己思考任务步骤,提升自动化效率

InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection

  • 分两阶段训练,先学看懂界面,再学会分步推理
  • 在多个测试中表现优异,尤其在复杂任务上超越传统方法
  • 适合需要智能操作界面的自动化场景,如软件测试

图形用户界面(GUI)代理借助多模态大语言模型,在计算机和手机等设备的任务自动化方面展现出巨大潜力。然而,现有代理在多步推理和依赖文本注释方面存在局限,影响了实际效果。我们提出 extit{InfiGUIAgent},一种基于多模态大语言模型的 GUI 代理,采用两阶段监督微调流程:第一阶段提升基础能力,如界面理解与定位;第二阶段通过合成数据引入层级推理与预期-反思推理机制,赋予代理原生推理能力。 extit{InfiGUIAgent} 在多个 GUI 基准测试中表现卓越,凸显原生推理对自动化交互的关键作用。相关资源已开源,地址见: exttt{https://github.com/Reallm-Labs/InfiGUIAgent}。

原文摘要 · Abstract (English)

Graphical User Interface (GUI) Agents, powered by multimodal large language models (MLLMs), have shown great potential for task automation on computing devices such as computers and mobile phones. However, existing agents face challenges in multi-step reasoning and reliance on textual annotations, limiting their effectiveness. We introduce \textit{InfiGUIAgent}, an MLLM-based GUI Agent trained with a two-stage supervised fine-tuning pipeline. Stage 1 enhances fundamental skills such as GUI understanding and grounding, while Stage 2 integrates hierarchical reasoning and expectation-reflection reasoning skills using synthesized data to enable native reasoning abilities of the agents. \textit{InfiGUIAgent} achieves competitive performance on several GUI benchmarks, highlighting the impact of native reasoning skills in enhancing GUI interaction for automation tasks. Resources are available at \url{https://github.com/Reallm-Labs/InfiGUIAgent}.

GUI代理多模态推理能力自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。