arXiv:2603.14448cs.LG2026-03被引 4

无需训练,通过逐步聚焦界面元素实现自然语言指令精准定位

Zoom to Essence: Trainless GUI Grounding by Inferring upon Interface Elements

  • 利用推理缩放技术逐步细化界面元素定位
  • 在多个基准上达到或超越现有最优水平
  • 适合追求零训练成本的GUI智能体开发者

基于多模态大模型的图形用户界面(GUI)智能体发展迅速,其核心能力是将自然语言指令映射到目标界面元素。现有方法通常需在大规模数据集上微调大模型,不仅带来高昂标注成本,且性能依赖数据质量与分布。我们发现复杂界面可分解为大模型直接理解的基础视觉元素,因此提出ZoomUI:通过推理缩放引导通用大模型逐步锚定指令目标元素。具体而言,先优化潜在思维,将原始指令转化为元素视觉特征描述;再利用内部注意力机制迭代聚焦目标界面区域。大量基准测试表明,ZoomUI达到甚至超越当前最优基线。

原文摘要 · Abstract (English)

Multimodal Large Language Model (MLLM)-based Graphical User Interface (GUI) agents develop rapidly, with visual grounding that maps natural language instructions to target UI elements serving as the core capability. Existing GUI agents typically fine-tune MLLM on massive datasets to handle challenges in understanding instructions and UI interfaces, which not only incurs high data annotation costs but also makes performance dependent on data quality and distribution. To avoid such cumbersome yet ineffective training, we notice that complex UI interfaces can be decomposed into basic visual elements directly understandable by common MLLMs. Consequently, we propose ZoomUI that leverages inference scaling to guide common MLLMs in progressively anchor instruction elements to increasingly detailed interface elements. Specifically, ZoomUI first optimizes the latent thinking to transform original instruction into element visual features description, and subsequently leverages internal attention to iteratively zoom in target element interface region. Evaluations on extensive benchmarks demonstrate that ZoomUI reaches or even surpasses SOTA baselines.

GUI智能体零训练推理缩放

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。