提出闭环式界面交互范式,让大模型更懂界面元素用法
What's Missing in Screen-to-Action? Towards a UI-in-the-Loop Paradigm for Multimodal GUI Reasoning

- 构建屏幕-元素-动作循环,显式学习界面元素定位与功能
- 在2.6万样本基准上实现当前最优界面理解效果
- 适合需要可解释性界面推理的AI助手开发场景
现有图形用户界面(GUI)推理任务仍具挑战性,尤其在界面理解方面。当前方法多依赖直接基于屏幕的决策,缺乏可解释性且未能全面理解界面元素,导致任务失败。为提升对界面的理解与交互能力,我们提出一种创新的GUI推理范式——界面在回路中(UI-in-the-Loop, UILoop)。该方法将GUI推理视为一个循环的屏幕-界面元素-动作过程,使多模态大语言模型(MLLMs)能够显式学习关键界面元素的定位、语义功能与实际用途,从而实现精准元素发现和可解释推理。此外,我们设计了一个聚焦于界面元素的更具挑战性的界面理解任务,并引入三项评估指标。相应地,我们构建了一个包含26,000个样本的基准数据集(UI Comprehension-Bench),用于全面评估现有方法对界面元素的掌握程度。大量实验表明,UILoop在界面理解性能上达到当前最优水平,并在GUI推理任务中取得更优结果。
原文摘要 · Abstract (English)
Existing Graphical User Interface (GUI) reasoning tasks remain challenging, particularly in UI understanding. Current methods typically rely on direct screen-based decision-making, which lacks interpretability and overlooks a comprehensive understanding of UI elements, ultimately leading to task failure. To enhance the understanding and interaction with UIs, we propose an innovative GUI reasoning paradigm called UI-in-the-Loop (UILoop). Our approach treats the GUI reasoning task as a cyclic Screen-UI elements-Action process. By enabling Multimodal Large Language Models (MLLMs) to explicitly learn the localization, semantic functions, and practical usage of key UI elements, UILoop achieves precise element discovery and performs interpretable reasoning. Furthermore, we introduce a more challenging UI Comprehension task centered on UI elements with three evaluation metrics. Correspondingly, we contribute a benchmark of 26K samples (UI Comprehension-Bench) to comprehensively evaluate existing methods' mastery of UI elements. Extensive experiments demonstrate that UILoop achieves state-of-the-art UI understanding performance while yielding superior results in GUI reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。