arXiv:2605.11218cs.AI2026-05

图像中的数字锚点会严重误导视觉语言模型的质量判断。

Don't Look at the Numbers: Visual Anchoring Bias and Layer-wise Representation in VLMs

论文配图:Don't Look at the Numbers: Visual Anchoring Bias and Layer-wise Representation in VLMs
图 1 · 摘自论文原文
  • 通过层级探针发现数字锚点在深层特征中饱和,但最佳质量预测在更深层。
  • 数字干扰的影响是图像质量严重退化的2.5倍,远超视觉变化影响。
  • 不同模型融合机制差异大,有的早期融合,有的几乎不融合。

图像中嵌入的数字锚点会系统性地误导六种来自五个架构族的视觉语言模型对图像质量的判断(方差分析η²=0.18-0.77,所有p<0.001)。锚点效应是严重图像质量退化影响的2.5倍,说明偏差并非源于视觉变化。分层探测显示:锚点分类达到饱和的层(L12-L34)反而不适合质量预测,最优预测层更靠后(决定系数R²=0.69-0.91)。融合分析揭示架构依赖的整合模式——两种模型在浅层(L1-L2)即完成融合,其余三种则部分或完全无融合。研究建立了视觉锚定偏差的因果解释,将行为易感性与表征动态联系起来。

原文摘要 · Abstract (English)

Embedded numeric anchors on images systematically bias Vision-Language Model quality judgments across six VLMs from five architectural families (ANOVA eta^2 = 0.18-0.77, all p < 0.001). Anchor effects are 2.5x larger than severe image quality degradation, confirming bias is not reducible to visual changes. Layer-wise probing reveals consistent dissociation: layers where anchor classification saturates (L12-L34) are suboptimal for quality prediction, with optimal layers deeper (R^2 = 0.69-0.91). Fusion analysis identifies architecture-dependent integration -- instant fusion at L1-L2 in two models versus partial or no fusion in three others. These results establish a causal account of visual anchoring bias, linking behavioral susceptibility to representation dynamics.

视觉语言模型认知偏差表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。