发现视觉语言模型的因果推理存在抽象鸿沟,真实推理能力远低于表面流畅度。
The Abstraction Gap in Vision-Language Causal Reasoning

- 设计双探针评估法,分离语言流畅性与真实因果链生成能力
- 7个模型在4.95万题上抽象鸿沟超0.5,文本得分6-8但链式推理不足2.5
- 仅一个模型接近零鸿沟,说明现有架构具备潜力但依赖预训练和结构
视觉语言模型(VLMs)能生成流畅的因果解释,但现有评估无法区分语言合理性与真实因果推理。本文提出双探针方法:文本探针衡量语言质量;链式文本探针要求模型先生成明确因果链。抽象鸿沟(AG)量化两者性能差值。在涵盖5,500张图像、49,500个问题的CAGE基准(覆盖Pearl因果层次)上评估8个VLM,发现7个模型的AG超过0.50,文本得分6–8,但链式得分低于2.5。在4.5万条链式标注数据上微调仍无法缩小差距。然而,有一个模型实现近乎零的AG。表明真实因果推理能力存在于当前架构中,取决于预训练和模型结构选择。CAGE为诊断VLM的忠实因果推理提供了工具。
原文摘要 · Abstract (English)
Vision-language models (VLMs) generate fluent causal explanations, but current evaluations cannot distinguish linguistic plausibility from faithful causal reasoning. We introduce a dual-probe methodology that isolates these properties. The Text-Only Probe measures linguistic quality. The Chain-Text Probe requires models to first generate explicit causal chains. The Abstraction Gap (AG) metric quantifies the normalized performance difference. Evaluating eight VLMs on CAGE (Causal Abstraction Gap Evaluation), a benchmark of 49,500 questions across 5,500 images spanning Pearl's causal hierarchy, we find seven models exhibit AG exceeding 0.50 with text scores of 6--8 but chain scores below 2.5. Fine-tuning on 45,000 chain-annotated examples fails to close the gap. However, one model achieves near-zero AG. The capability exists within current VLM architectures and depends on pretraining and architectural choices. CAGE provides a diagnostic tool for assessing faithful causal reasoning in VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。