用灰图替换测试,发现模型空间推理的解码能力不等于真实视觉感知。
Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning
- 用灰图作为因果控制,揭示模型解码与实际感知的差异。
- 水平方向真实依赖图像,垂直方向是先验,深度方向则符号反转。
- 该方法可自校准,适合验证视觉语言模型的隐含知识真实性。
通过将图像替换为灰色空白,我们发现标准的线性探针和转向恢复无法准确反映视觉语言模型(VLM)在图像中的真实视觉接地。在空间推理任务中,探针在各轴上的解码准确率达73%–97%,训练无关的投影将接近随机的轴从59%提升至79%,看似解锁了隐含知识,但灰图仲裁者揭示了三种被探针混淆的接地状态:真实接地(视觉依赖、正确)、先验(视觉独立、仅方向默认)、以及意外反转:虽可解码且因果可控,但符号错误导致性能低于随机水平。这一分类在14个跨6种语言模型家族(2B–27B参数量)中一致成立:水平方向为接地,垂直方向为先验,深度方向为反转,且反转现象随模型规模在家族内显现。八模型中七例复现解码与部署符号反转,最小修正方式因几何结构而异:干净模型只需无训练旋转,复杂反转需训练低秩编辑,揭示出模型修正复杂度谱。该低成本自校准仲裁器能清晰区分真实感知、反向感知与先验替代,建议成为验证VLM隐含知识与转向声明的默认控制手段。
原文摘要 · Abstract (English)
The standard way to read latent knowledge out of a model, a linear probe confirmed by a steering recovery, can systematically overstate what a vision-language model (VLM) actually grounds in the image. We show this on spatial reasoning, where the error is invisible to both probing and steering yet exposed by a one-line causal control: replacing the image with a gray blank. Probes decode the within-axis answer at 73--97% across axes, and a training-free projection lifts a near-chance axis from 59% to 79%, exactly the signature of unlocking latent knowledge. The blank-image arbiter refutes it, revealing three grounding regimes that probing conflates: an axis can be grounded (vision-dependent, correct), a prior (vision-independent, with its decode and its apparent recovery a directional default rather than perception), or, surprisingly, inverted: decodable, causally controllable, but deployed with the wrong sign, so the model scores below chance and the error requires looking. The taxonomy holds across the studied VLMs: in fourteen models spanning six language-model families and 2B--27B, horizontal is grounded, vertical is a prior, and depth is inverted, with the inversion emerging at scale within families. The decode-versus-deploy inversion replicates on seven of eight models across five families, and the minimal edit that re-deploys it varies with geometry: a training-free rotation matches a trained edit on the cleanest model, while distributed inversions need a trained low-rank edit, tracing a per-model correction-complexity spectrum. The cheap, self-calibrating arbiter cleanly separates grounded perception, inverted perception, and prior substitution; we argue it should be a default control for latent-knowledge and steering claims in VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。