arXiv:2512.09396cs.MAcs.AI2025-12

用多模型协作提升GUI自动化智能水平

GAIR: GUI Automation via Information-Joint Reasoning and Group Reflection

  • 通过联合推理融合多个专用模型能力
  • 在基准测试中显著提升任务完成率
  • 适合需要跨场景自动化系统的研究者

构建用于GUI自动化的AI系统已吸引大量研究关注,其中多模态大模型(MLLMs)被用于处理用户需求并生成操作。然而,GUI自动化涵盖从文档处理到在线购物、从CAD到视频编辑的广泛任务,不同任务间的差异要求模型具备异构能力与多维专长,带来建模挑战。为此,我们提出GAIR:基于信息联合推理与群体反思的GUI自动化框架,该框架通过集成异构模型的知识与能力,构建高性能的自动化代理系统。由于不同特定于GUI的MLLMs在不同数据集上训练,各有优势,GAIR引入一个通用型MLLM,联合处理来自多个特定模型的信息,从而提升系统性能。该通用模型同时担任决策者,基于已有信息尝试执行合理操作。当其判断信息不足时,系统进入群体反思状态:通用模型根据各特定模型的优劣势提供差异化指令与提示,驱动它们更精准地收集关键信息,支持深层推理与决策。我们在多个GUI基准上进行了广泛实验,验证了GAIR的有效性与可靠性。

原文摘要 · Abstract (English)

Building AI systems for GUI automation task has attracted remarkable research efforts, where MLLMs are leveraged for processing user requirements and give operations. However, GUI automation includes a wide range of tasks, from document processing to online shopping, from CAD to video editing. Diversity between particular tasks requires MLLMs for GUI automation to have heterogeneous capabilities and master multidimensional expertise, raising problems on constructing such a model. To address such challenge, we propose GAIR: GUI Automation via Information-Joint Reasoning and Group Reflection, a novel MLLM-based GUI automation agent framework designed for integrating knowledge and combining capabilities from heterogeneous models to build GUI automation agent systems with higher performance. Since different GUI-specific MLLMs are trained on different dataset and thus have different strengths, GAIR introduced a general-purpose MLLM for jointly processing the information from multiple GUI-specific models, further enhancing performance of the agent framework. The general-purpose MLLM also serves as decision maker, trying to execute a reasonable operation based on previously gathered information. When the general-purpose model thinks that there isn't sufficient information for a reasonable decision, GAIR would transit into group reflection status, where the general-purpose model would provide GUI-specific models with different instructions and hints based on their strengths and weaknesses, driving them to gather information with more significance and accuracy that can support deeper reasoning and decision. We evaluated the effectiveness and reliability of GAIR through extensive experiments on GUI benchmarks.

GUI自动化多模型协作决策增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。