arXiv:2607.26913cs.CV2026-07

旧语言描述会误导视觉判断,关键在于模型如何组织变化方向。

Prior Directions: Why GUI Grounding Gets Locked in the Past

论文配图:Prior Directions: Why GUI Grounding Gets Locked in the Past
图 1 · 摘自论文原文
  • 发现模型受旧描述影响时,变化集中在少数固定方向上。
  • 移除这些方向可恢复正确判断,移除其他部分无效。
  • 适合研究视觉语言模型偏差与鲁棒性的人阅读。

视觉语言模型常依赖先前视觉状态的描述来决策当前场景。当场景改变时,过时的语言描述可能将原本正确的视觉判断引向错误答案。我们在此控制环境下研究这一失效现象——仅改变语言先验。实验发现,锁入程度更强的模型在最终答案前的表征变化更小。这表明锁入不取决于表征移动距离,而取决于移动方式。在难以纠正的模型中,先验引发的变化集中于一组紧凑的方向,这些方向在不同样本中反复出现,我们称之为‘先验方向’(Prior Directions)。它们在保留样本中持续重现,且四模型对比显示方向集中度越高,锁入越强。可控干预证明:移除与先验方向对齐的分量可恢复视觉定位能力,而移除同等大小的正交分量则影响甚微。因此,先验控制的关键在于先验引发的变化是否形成一致且可复用的表征模式。该机制解释了为何同一先验在一个模型中可被修正,而在另一模型中却占据主导。

原文摘要 · Abstract (English)

Vision-language models often use descriptions of earlier visual states to make decisions about the current scene. When the scene changes, stale language can redirect an otherwise correct visual judgment toward an outdated answer. We study this failure as visual lock-in in a controlled grounding setting where only the verbalized prior varies. Across models, stronger lock-in accompanies smaller changes in the model representation before the final answer. This reversal suggests that lock-in depends not on how far this representation moves, but on how that movement is organized. In models that are harder to correct, prior-induced changes concentrate along a compact set of directions that repeatedly appear across examples. We call these recurrent axes the Prior Directions. They recur on held-out examples, while a descriptive four-model comparison associates greater concentration with stronger lock-in. Controlled interventions show that removing the component aligned with the Prior Directions restores visual grounding, whereas removing an equally large orthogonal component has little effect. Prior control thus arises when prior-induced changes form a coherent and reusable pattern in the representation used to produce the answer. This account explains why the same prior remains revisable in one model yet becomes dominant in another.

视觉语言模型偏差表征方向鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。