arXiv:2508.06220cs.CLcs.AI2025-08被引 1

测试视觉语言模型能否从图文信息中进行因果推理。

InfoCausalQA:Can Models Perform Non-explicit Causal Reasoning Based on Infographic?

  • 构建图文结合的因果推理基准,区分数值与语义两类推理任务。
  • 当前模型在两类推理上表现远低于人类,尤其在语义因果上差距明显。
  • 适合关注多模态认知能力、因果推理的AI研究者参考。

视觉语言模型(VLMs)在感知与推理方面取得显著进展,但其在因果推断——人类认知的核心能力——方面的表现仍不充分,尤其是在多模态场景中。本文提出InfoCausalQA,一个基于图文信息的新型基准,用于评估模型在结构化视觉数据与文本上下文结合下的因果推理能力。该基准包含两项任务:任务1聚焦于基于推断数值趋势的量化因果推理;任务2针对五类语义因果关系(原因、结果、干预、反事实、时间顺序)的推理。研究从四个公开来源手动收集了494对图文数据,并使用GPT-4o生成1,482个高质量的多选问答对,经人工审核确保问题无法仅通过表面线索回答,而需真正实现视觉锚定。实验结果表明,当前VLMs在计算推理上表现有限,在语义因果推理上更是存在显著不足,其性能远低于人类水平,揭示出多模态系统在利用图文信息进行因果推断方面仍有巨大提升空间。通过InfoCausalQA,我们强调了推动多模态AI因果推理能力发展的必要性。

原文摘要 · Abstract (English)

Recent advances in Vision-Language Models (VLMs) have demonstrated impressive capabilities in perception and reasoning. However, the ability to perform causal inference -- a core aspect of human cognition -- remains underexplored, particularly in multimodal settings. In this study, we introduce InfoCausalQA, a novel benchmark designed to evaluate causal reasoning grounded in infographics that combine structured visual data with textual context. The benchmark comprises two tasks: Task 1 focuses on quantitative causal reasoning based on inferred numerical trends, while Task 2 targets semantic causal reasoning involving five types of causal relations: cause, effect, intervention, counterfactual, and temporal. We manually collected 494 infographic-text pairs from four public sources and used GPT-4o to generate 1,482 high-quality multiple-choice QA pairs. These questions were then carefully revised by humans to ensure they cannot be answered based on surface-level cues alone but instead require genuine visual grounding. Our experimental results reveal that current VLMs exhibit limited capability in computational reasoning and even more pronounced limitations in semantic causal reasoning. Their significantly lower performance compared to humans indicates a substantial gap in leveraging infographic-based information for causal inference. Through InfoCausalQA, we highlight the need for advancing the causal reasoning abilities of multimodal AI systems.

因果推理多模态视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。