arXiv:2508.16271cs.CVcs.LG2025-08

用新训练方法提升AI理解界面坐标的准确率

Structuring GUI Elements through Vision Language Models: Towards Action Space Generation

  • 设计基于交并比的坐标采样数据增强策略
  • 在UI坐标生成任务上显著优于传统训练方法
  • 适合做智能交互、自动化测试的研究者参考

多模态大语言模型(MLLM)在人机交互中扮演关键角色。本文聚焦于利用MLLM进行图形用户界面(GUI)元素结构化,根据屏幕内容处理用户指令。尽管前景广阔,但其在精确生成界面元素坐标方面表现受限,原因在于语言模型训练中数值坐标的语义空洞。为此,我们提出一种交并比增强的最大似然(IAML)训练范式。该方法通过基于交并比的坐标采样构建增强数据集,考虑与真实坐标的接近程度,并在此基础上微调MLLM,以缓解传统最大似然估计中的暴露偏差问题。大量实验表明,IAML训练方法在坐标生成任务上显著优于传统范式。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have emerged as pivotal tools in enhancing human-computer interaction. In this paper we focus on the application of MLLMs in the field of graphical user interface (GUI) elements structuring, where they assist in processing user instructions based on screen contents. Despite the promise of MLLMs, their performance in precisely generating UI element coordinates, a critical aspect of GUI understanding, is hindered by the nature of next-token prediction training. This challenge arises from the semantic void surrounding numerical UI coordinates in language representation spaces, necessitating a substantial and diverse dataset to bolster visual module capabilities. To address these limitations, we introduce an IoU-Augmented Maximum Likelihood (IAML) training paradigm. Specifically, our approach involves a novel pipeline for IoU-based coordinate sampling to augment the training data, which considers the proximity to ground truth coordinates. This data augmentation strategy is then employed to fine-tune MLLMs under the IAML paradigm, which is designed to mitigate the exposure bias problem inherent in traditional maximum likelihood estimation. Through extensive experiments, we demonstrate the superior performance of our IAML training approach over traditional training paradigms.

GUI理解多模态模型坐标生成训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。