研究视觉语言模型在路径追踪中的失败原因,发现其易受局部相似干扰。
VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following

- 设计可控路径追踪任务,隔离局部竞争问题
- 顶级模型在相似干扰下频繁偏离目标路径
- 适合关注多模态模型可靠性与鲁棒性的人阅读
视觉语言模型(VLMs)在多模态基准上表现优异,但在基础视觉操作上仍缺乏稳健控制。本文研究了‘线段追踪’任务,即模型需通过连续局部延续跟踪选定视觉路径。为隔离该能力,设计了受控追踪任务,引入附近竞争路径,同时降低语义和拓扑模糊性(如交叉与重叠)。在这些任务中,即使最先进的VLMs也频繁丢失目标路径并切换到邻近替代路径,尤其当这些替代路径在局部外观上与目标相似时。行为干预与内部分析表明,失败源于局部竞争:邻近的相似干扰项会将模型拉离真实延续路径。标准缓解方法无效:模型规模扩展仅带来有限改善,推理部分通过高成本替代策略补偿,而显式追踪指令无法恢复稳定路径跟随。最终,在更复杂的缠绕电缆场景和地铁地图测试中,相同的路径切换错误仍持续存在。
原文摘要 · Abstract (English)
Vision-language models (VLMs) achieve strong performance on multimodal benchmarks, but may still lack robust control over basic visual operations. We study \textit{line tracing}, where a model must follow a selected visual path through successive local continuations. To isolate this ability, we design controlled tracing tasks that introduce nearby competitors while reducing semantic and topological ambiguity such as crossings and overlaps. Across these tasks, even state-of-the-art VLMs frequently lose the target path and switch to nearby alternatives, especially when those alternatives look locally similar to the target. Behavioral interventions and internal analyses indicate that these failures arise from local competition: nearby similar distractors pull the model away from the true continuation. Standard remedies do not remove this bottleneck: model-size scaling provides only limited gains, reasoning partially compensates through costly substitute strategies, and explicit tracing instructions fail to recover stable path following. Finally, tests on tangled-cable scenes and metro maps with richer visual complexity show that the same path-switching failure persists beyond our controlled settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。