视觉参考下,大模型回忆事实知识能力下降一半。
Can VLMs Recall Factual Associations From Visual References?
- 用视觉而非文字提示时,模型知识回忆能力减半。
- 内部状态探测可92%准确识别不可靠回答场景。
- 无需重训练,能提升视觉问答任务覆盖与准确率。
通过控制实验,我们发现视觉语言模型(VLMs)在多模态对齐方面存在系统性缺陷。当提供实体的文本参考时,VLMs 能有效回忆事实关联;但若参考为图像,则其回忆能力显著下降。强制模型依赖实体的图像表示,使其知识回忆能力降低约50%,表明模型难以将内部知识与图像表征关联。我们发现此类关联失败与模型内部状态的特定模式相关,基于这些状态的探测器在识别模型不可靠响应时准确率达92%以上。该探测器无需重新训练即可用于识别需要多模态理解的问题中的潜在错误。应用于视觉问答任务时,该方法使覆盖范围提升7.87%(绝对值),同时将错误风险降低0.9%(绝对值)。解决这一可检测的系统性缺陷是语言对齐研究的重要方向,并为此提出未来改进建议。
原文摘要 · Abstract (English)
Through a controlled study, we identify a systematic deficiency in the multimodal grounding of Vision Language Models (VLMs). While VLMs can recall factual associations when provided a textual reference to an entity; their ability to do so is significantly diminished when the reference is visual instead. Forcing VLMs to rely on image representations of an entity halves their ability to recall factual knowledge, suggesting that VLMs struggle to link their internal knowledge of an entity with its image representation. We show that such linking failures are correlated with the expression of distinct patterns in model internal states, and that probes on these internal states achieve over 92% accuracy at flagging cases where the VLM response is unreliable. These probes can be applied, without retraining, to identify when a VLM will fail to correctly answer a question that requires an understanding of multimodal input. When used to facilitate selective prediction on a visual question answering task, the probes increase coverage by 7.87% (absolute) while also reducing the risk of error by 0.9% (absolute). Addressing the systematic, detectable deficiency is an important avenue in language grounding, and we provide informed recommendations for future directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。