用结构化图谱诊断视觉语言模型的因果推理缺陷。
Diagnosing Causal Reasoning in Vision-Language Models via Structured Relevance Graphs
- 构建查询相关的因果图,显式编码物体、属性和关系
- 新基准测试显示,结构化信息提升因果归因与推理一致性
- 适合研究模型可解释性与因果推理的学者
大型视觉语言模型(LVLMs)在视觉问答任务中表现优异,但常依赖虚假相关而非真实因果推理。现有评估仅关注答案正确性,无法区分失败是源于推理能力不足还是因果信息识别错误。本文提出视觉语言因果图(VLCGs),一种查询条件下的结构化表示,显式编码因果相关的物体、属性、关系及场景假设。基于此,构建了ViLCaR诊断基准,包含因果归因、因果推断和问答任务,并设计与图对齐的评估指标,超越最终答案准确率,评估相关性识别能力。在主流LVLM上的实验表明,注入结构化相关性信息显著提升了归因与推断一致性,优于零样本和标准上下文学习。结果表明,当前LVLM因果推理限制主要源于缺乏结构引导,而非推理能力不足。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) achieve strong performance on visual question answering benchmarks, yet often rely on spurious correlations rather than genuine causal reasoning. Existing evaluations primarily assess the correctness of the answers, making it unclear whether failures arise from limited reasoning capability or from misidentifying causally relevant information. We introduce Vision-Language Causal Graphs (VLCGs), a structured, query-conditioned representation that explicitly encodes causally relevant objects, attributes, relations, and scene-grounded assumptions. Building on this representation, we present ViLCaR, a diagnostic benchmark comprising tasks for Causal Attribution, Causal Inference, and Question Answering, along with graph-aligned evaluation metrics that assess relevance identification beyond final answer accuracy. Experiments in state-of-the-art LVLMs show that injecting structured relevance information significantly improves attribution and inference consistency compared to zero-shot and standard in-context learning. These findings suggest that current limitations in LVLM causal reasoning stem primarily from insufficient structural guidance rather than a lack of reasoning capacity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。