视觉语言模型在纹理倾斜感知上存在锚点式错误,无法准确表达连续几何信息。
Anchored, Not Graded: Vision-Language Models Fail at Slant-from-Texture Perception

- 模型仅在特定角度(如0°、±25°、±45°)预测倾斜,缺乏连续响应
- 无论视场、光学倾斜或曲率如何变化,预测结果均不敏感
- 微调可部分改善,但锚点偏差仍持续存在,提示接口设计缺陷
人类通过纹理感知表面倾斜时表现出系统性、渐变的偏差,这在心理物理学实验中稳定出现。先前研究显示无监督CNN能复现部分人类类似偏差,而有监督CNN则不能。本研究考察视觉语言模型(VLMs)是否具备此类能力。在多个VLM家族及不同规模下,零样本和上下文提示均导致显著失败:倾斜预测仅集中在少数锚点角度(如0°、±25°、±45°),对刺激视场、光学倾斜或表面曲率几乎无依赖。有监督微调虽部分缓解该问题,但残余锚定现象依然存在。我们认为,这种锚定反映的是表示到输出语言接口的失败——并非缺乏几何编码,而是无法以渐变形式表达几何信息。尽管高阶视觉语言任务可能不依赖低层几何线索,此现象揭示了模型在精确几何理解上的根本局限。
原文摘要 · Abstract (English)
Human perception of surface slant from texture exhibits systematic, graded biases that emerge reliably in psychophysical experiments. Prior work showed that unsupervised CNNs reproduce several human-like biases, while supervised CNNs do not. Do Vision-Language Models (VLMs) exhibit similar competences? Across multiple VLM families and model scales, zero-shot and in-context prompting both produce distinctive failures: slant is predicted at only a small set of anchors (e.g., 0\degree, $\pm$25\degree, $\pm$45\degree) with little dependence on stimulus field of view, optical slant, or surface curvature. Supervised fine-tuning partially remediates the failure, but residual anchoring persists. While success in high-level vision-language benchmarks might not require sensitivity to low-level geometric cues, we interpret anchoring as a failure at the representation-to-output language interface: not necessarily an absence of geometric encoding, but a failure to express it in a graded form.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。