通过视觉反事实测试,揭示视觉模型如何权衡先验知识与图像信息。
Pixels Versus Priors: Controlling Knowledge Priors in Vision-Language Models through Visual Counterfacts
- 构建视觉反事实数据集,让图像与常识形成冲突
- 模型中后期层更依赖图像而非记忆中的先验知识
- 提出激活干预方法,可精准控制模型偏向视觉或知识
多模态大语言模型在视觉问答等任务上表现优异,但其推理机制究竟是依赖于记忆中的世界知识,还是输入图像中的视觉信息仍不明确。为此,我们提出视觉反事实(Visual CounterFact)数据集,包含一系列视觉逼真的反事实样本,使常识先验(如红草莓)与图像输入(如蓝草莓)直接冲突。实验表明,模型初始预测受记忆先验主导,但在中后期层逐渐转向视觉证据。这一动态过程揭示了两种模态间的竞争关系,最终视觉输入在评估阶段覆盖先验。为调控此行为,我们提出像素对抗先验(Pixels Versus Priors, PvP)引导向量,通过激活层面干预实现输出控制。平均而言,PvP成功将99.3%的颜色预测和80.8%的尺寸预测从先验转向反事实。这些成果为解析与调控多模态模型的事实性行为提供了新工具。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) perform well on tasks such as visual question answering, but it remains unclear whether their reasoning relies more on memorized world knowledge or on the visual information present in the input image. To investigate this, we introduce Visual CounterFact, a new dataset of visually-realistic counterfactuals that put world knowledge priors (e.g, red strawberry) into direct conflict with visual input (e.g, blue strawberry). Using Visual CounterFact, we show that model predictions initially reflect memorized priors, but shift toward visual evidence in mid-to-late layers. This dynamic reveals a competition between the two modalities, with visual input ultimately overriding priors during evaluation. To control this behavior, we propose Pixels Versus Priors (PvP) steering vectors, a mechanism for controlling model outputs toward either world knowledge or visual input through activation-level interventions. On average, PvP successfully shifts 99.3% of color and 80.8% of size predictions from priors to counterfactuals. Together, these findings offer new tools for interpreting and controlling factual behavior in multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。