arXiv:2503.14021cs.CVcs.AI2025-03CVPR被引 19

用多模态模型提升GUI理解,解决空间结构建模难题

MP-GUI: Modality Perception with MLLMs for GUI Understanding

  • 设计三模态感知器提取图形、文本、空间信息
  • 融合门控机制自适应组合,适配不同任务需求
  • 自动构建数据集,缓解标注数据稀缺问题

图形用户界面(GUI)在现代社会中日益重要,理解其语义对以人为本的系统至关重要。然而,与自然图像或文档不同,GUI由人工设计的图形元素组成,需通过特定布局传达语义。当前多模态大语言模型虽能处理图文内容,但在缺乏显式空间结构建模的情况下,仍难以准确理解GUI。此外,由于隐私问题和环境噪声,高质量空间结构数据获取困难。为此,我们提出MP-GUI,一种专为GUI理解设计的多模态大语言模型。该模型包含三个专用感知器,分别提取图形、文本和空间模态特征,并通过空间结构优化策略与自适应融合门控机制,结合生成符合不同任务偏好的视觉线索。针对训练数据稀缺问题,我们还设计了自动化数据采集流程。大量实验表明,即使在有限数据条件下,MP-GUI在多种GUI理解任务上仍表现优异。

原文摘要 · Abstract (English)

Graphical user interface (GUI) has become integral to modern society, making it crucial to be understood for human-centric systems. However, unlike natural images or documents, GUIs comprise artificially designed graphical elements arranged to convey specific semantic meanings. Current multi-modal large language models (MLLMs) already proficient in processing graphical and textual components suffer from hurdles in GUI understanding due to the lack of explicit spatial structure modeling. Moreover, obtaining high-quality spatial structure data is challenging due to privacy issues and noisy environments. To address these challenges, we present MP-GUI, a specially designed MLLM for GUI understanding. MP-GUI features three precisely specialized perceivers to extract graphical, textual, and spatial modalities from the screen as GUI-tailored visual clues, with spatial structure refinement strategy and adaptively combined via a fusion gate to meet the specific preferences of different GUI understanding tasks. To cope with the scarcity of training data, we also introduce a pipeline for automatically data collecting. Extensive experiments demonstrate that MP-GUI achieves impressive results on various GUI understanding tasks with limited data.

GUI理解多模态模型空间结构自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。