arXiv:2506.00869cs.CL2025-06Conference of the …被引 1

新基准揭示视觉语言模型因果推理能力严重不足

What's Missing in Vision-Language Models? Probing Their Struggles with Causal Order Reasoning

  • 设计两个新基准,专攻因果顺序推理,排除表面线索干扰
  • 模型在因果任务上仅略高于随机猜测,远低于识别能力
  • 训练数据缺乏因果表达是主因,针对性微调可提升表现

尽管视觉语言模型在下游任务中表现优异,但其对视觉输入中因果关系的理解与推理能力仍不明确。稳健的因果推理是解决复杂高层推理任务的基础,然而现有评估基准常混合多种推理类型,模型可依赖物体识别和活动判断作为捷径获得正确答案,难以真实评估其因果推理能力。为此,我们提出VQA-Causal和VCR-Causal两个新基准,专门用于隔离并严格评估模型的因果推理能力。研究发现,虽然模型在物体和活动识别上表现优异,但在因果推理任务中表现不佳,往往仅略高于随机猜测。进一步分析表明,这一局限源于广泛使用训练数据中因果表达严重缺失,因果关系极少被明确传达。我们还探索了基于困难负样本的微调策略,结果显示针对性微调可提升模型因果推理能力,同时保持泛化性和下游性能。本研究揭示了当前视觉语言模型的关键短板,并为未来因果理解研究奠定基础。

原文摘要 · Abstract (English)

Despite the impressive performance of vision-language models (VLMs) on downstream tasks, their ability to understand and reason about causal relationships in visual inputs remains unclear. Robust causal reasoning is fundamental to solving complex high-level reasoning tasks, yet existing benchmarks often include a mixture of reasoning questions, and VLMs can frequently exploit object recognition and activity identification as shortcuts to arrive at the correct answers, making it challenging to truly assess their causal reasoning abilities. To bridge this gap, we introduce VQA-Causal and VCR-Causal, two new benchmarks specifically designed to isolate and rigorously evaluate VLMs' causal reasoning abilities. Our findings reveal that while VLMs excel in object and activity recognition, they perform poorly on causal reasoning tasks, often only marginally surpassing random guessing. Further analysis suggests that this limitation stems from a severe lack of causal expressions in widely used training datasets, where causal relationships are rarely explicitly conveyed. We additionally explore fine-tuning strategies with hard negative cases, showing that targeted fine-tuning can improve model's causal reasoning while maintaining generalization and downstream performance. Our study highlights a key gap in current VLMs and lays the groundwork for future work on causal understanding.

视觉语言模型因果推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。