揭示视觉语言模型在模糊图像描述中的认知偏见与决策机制
Vision-Language Asymmetry in Bistable Image Captioning

- 通过83张双义图像测试,发现模型存在默认主导、强制主导和平衡三种响应模式
- 72%的图像在视觉层同时激活双解释特征,但语言层仍呈现单向选择
- 视觉表征的并存不等于语言输出的可调,决策瓶颈在视觉之后
Wittgenstein的鸭兔图引发一个问题:当模型对模糊图像生成描述时,对某一解释的承诺究竟在模型何处做出?我们基于83个双义刺激,进行3,320次生成行为基准测试,揭示了中性提示与强制选择提示下的三种响应模式(默认主导、强制主导、强制平衡)。利用在LLaVA-1.6-7B实际使用的CLIP层上训练的TopK稀疏自编码器(验证误差0.93),对69个具备双视角特征池的刺激进行探测。结果显示,72%(50/69)的图像在视觉塔层同时激活两个特征池,包括全部12个默认主导的鸭兔图及8个强制平衡的年轻/年长图中的7个。在CLIP层22进行因果操控可使默认主导图的描述翻转(流畅性约束下兔子翻转率达33%),但在任何系数下均无法翻转强制平衡的年轻/年长图,尽管其视觉侧存在叠加态。说明主导性瓶颈位于视觉塔下游;视觉表征与语言输出之间的差距,为理解‘看见’与‘视为’提供了实证接口。另指出方法论问题:使用排名统计时需修正平局,否则会引入隐性行序偏差。
原文摘要 · Abstract (English)
Wittgenstein's duck-rabbit poses a question for vision-language models: when a model captions an ambiguous image, where in the model is the commitment to one aspect made? We address this with a 3,320-generation behavioral baseline over 83 bistable stimuli that surfaces three regimes (default-dominant, force-dominant, force-balanced) under neutral vs forced-choice prompting, then probe the underlying representations using a TopK sparse autoencoder we train on the CLIP layer that LLaVA-1.6-7B actually consumes (validation EV 0.93). Across 69 bistable stimuli with both per-aspect feature pools available, 72% (50/69) show simultaneous activation of both pools at the vision tower, including 12/12 default-dominant duck/rabbit and 7/8 force-balanced young/old. Causal steering at CLIP layer 22 flips captions on default-dominant stimuli (33% rabbit-flip rate under a fluency guard) but cannot flip captions on force-balanced young/old at any tested coefficient, despite their vision-side superposition. The dominance bottleneck lives downstream of the vision tower; the gap between vision-side representation and language-side commitment is an empirical handle on the seeing/seeing-as distinction. We also flag a methodological note: rank-based statistics on TopK SAE outputs require tie-corrected ranking to avoid silent row-order bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。