arXiv:2602.06351cs.AIcs.CV2026-02被引 3

通过融合文本与图标语义提升GUI定位准确率

Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal Fusion

  • 引入注意力、OCR文本和图标描述三重信号融合机制
  • 无需微调,在4个基准上均达领先性能
  • 适合需要低数据依赖的GUI智能体开发场景

GUI grounding 将自然语言指令映射到正确界面元素,是GUI代理的感知基础。现有方法主要依赖大规模GUI数据集微调多模态大语言模型(MLLMs)预测目标坐标,存在数据消耗大且泛化能力差的问题。近期基于注意力的方法虽无需任务微调,但因GUI图像缺乏显式互补空间锚点,定位可靠性不足。为此,我们提出Trifuse框架,通过共识-单峰(CS)融合策略,显式整合注意力、OCR提取的文本线索和图标级描述语义,强化跨模态一致性并保留锐利定位峰值。在四个基准上的大量评估表明,Trifuse无需任务微调即实现优异性能,显著降低对昂贵标注数据的依赖。消融实验进一步验证,引入OCR与描述线索能持续提升不同骨干网络的定位表现,证明其作为通用GUI定位框架的有效性。

原文摘要 · Abstract (English)

GUI grounding maps natural language instructions to the correct interface elements, serving as the perception foundation for GUI agents. Existing approaches predominantly rely on fine-tuning multimodal large language models (MLLMs) using large-scale GUI datasets to predict target element coordinates, which is data-intensive and generalizes poorly to unseen interfaces. Recent attention-based alternatives exploit localization signals in MLLMs attention mechanisms without task-specific fine-tuning, but suffer from low reliability due to the lack of explicit and complementary spatial anchors in GUI images. To address this limitation, we propose Trifuse, an attention-based grounding framework that explicitly integrates complementary spatial anchors. Trifuse integrates attention, OCR-derived textual cues, and icon-level caption semantics via a Consensus-SinglePeak (CS) fusion strategy that enforces cross-modal agreement while retaining sharp localization peaks. Extensive evaluations on four grounding benchmarks demonstrate that Trifuse achieves strong performance without task-specific fine-tuning, substantially reducing the reliance on expensive annotated data. Moreover, ablation studies reveal that incorporating OCR and caption cues consistently improves attention-based grounding performance across different backbones, highlighting its effectiveness as a general framework for GUI grounding.

GUI定位多模态融合注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。