通过迭代推理与反馈提升界面指令定位准确率
Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
- 不依赖训练,用多步推理逐步优化目标定位
- 在ScreenSpot Pro上达到68.4%准确率,提升4.8点
- 适用于真实工业界面,对模糊遮挡场景鲁棒
GUI定位旨在将自然语言指令与复杂用户界面中的精确区域对齐。先进多模态大模型虽具备较强视觉定位能力,但在小目标、视觉相似目标及真实布局模糊性上仍表现不佳,根源在于定位能力有限且未充分挖掘现有推理潜力。本文提出无需训练的多步定位框架Chain of Ground(CoG),利用多模态大模型进行迭代视觉推理与修正。模型不直接预测,而是逐步反思并调整假设,实现更精准、可解释的定位。在ScreenSpot Pro基准上,准确率达到68.4%,提升4.8点;为评估真实世界泛化能力,我们引入TPanel UI数据集,包含420个带模糊、遮挡等视觉失真的工业控制面板标注数据。在该数据集上,Chain of Ground相较强基线Qwen3 VL 235B提升6.9点,证明了多步无训练定位在真实与数字界面中的有效性。结果表明,通过结构化迭代精炼可释放潜在定位能力,而非依赖额外训练。
原文摘要 · Abstract (English)
GUI grounding aims to align natural language instructions with precise regions in complex user interfaces. Advanced multimodal large language models show strong ability in visual GUI grounding but still struggle with small or visually similar targets and ambiguity in real world layouts. These limitations arise from limited grounding capacity and from underuse of existing reasoning potential. We present Chain of Ground CoG a training free multi step grounding framework that uses multimodal large language models for iterative visual reasoning and refinement. Instead of direct prediction the model progressively reflects and adjusts its hypotheses leading to more accurate and interpretable localization. Our approach achieves 68.4 accuracy on the ScreenSpot Pro benchmark an improvement of 4.8 points. To measure real world generalization we introduce TPanel UI a dataset of 420 labeled industrial control panels with visual distortions such as blur and masking. On TPanel UI Chain of Ground improves over the strong baseline Qwen3 VL 235B by 6.9 points showing the effectiveness of multi step training free grounding across real world and digital interfaces. These results highlight a direction for unlocking grounding potential through structured iterative refinement instead of additional training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。