构建因果物理推理基准,让视觉语言模型学会真正理解物体间的因果关系。
Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs

- 设计4大领域3000+题目的因果图标注数据集,精准刻画物理关系
- 提出因果对齐度量指标,发现主流模型普遍遗漏关键因果链
- 引入因果推理微调方法,显著提升模型准确率与可解释性
理解物理世界是智能行为的基础,但当前先进视觉语言模型(VLMs)在因果物理推理上仍表现不佳,常给出看似合理却错误的答案。为弥补这一差距,我们提出CausalPhys,一个包含3000余道精心设计的视频与图像问题的基准,覆盖感知、预测、干预和目标导向四个领域。每道题均配有专家标注的因果图,刻画物体属性与事件之间的依赖关系,支持可解释的细粒度评估。基于此,我们构建了基于因果图的度量标准,定量评估模型思维链与正确因果关系的一致性,突破仅依赖答案准确率的局限,实现对因果推理失败的系统诊断。通过该指标分析主流VLMs,发现其在捕捉因果依赖方面存在系统性缺陷,亟需引入因果意识学习。为此,我们进一步提出因果推理引导微调(CRFT),显式对齐模型推理与因果结构。大量实验表明,CRFT在多个模型架构上均显著提升推理准确率与可解释性。CausalPhys通过统一数据集构建、因果评估与因果导向训练,为推动现代VLMs向因果化物理理解迈进奠定坚实基础。
原文摘要 · Abstract (English)
Understanding and reasoning about the physical world is the foundation of intelligent behavior, yet state-of-the-art vision-language models (VLMs) still fail at causal physical reasoning, often producing plausible but incorrect answers. To address this gap, we introduce CausalPhys, a benchmark of over 3,000 carefully curated video- and image-based questions spanning four domains: Perception, Anticipation, Intervention, and Goal Orientation. Each question is paired with an expert-annotated causal graph capturing object-attribute-event dependencies, enabling interpretable and fine-grained evaluation of causal understanding. Building on this, we formulate a causal-graph-grounded metric that quantitatively measures how well a model's chain-of-thought reasoning aligns with the correct causal relations, moving beyond answer-only accuracy and enabling systematic diagnosis of VLMs' causal reasoning failures. Using this metric, we conduct a comprehensive analysis of leading VLMs, revealing systematic gaps in capturing causal dependencies and underscoring the need for causality-aware learning. To address these limitations, we further propose Causal Rationale-informed Fine-Tuning (CRFT), which explicitly aligns VLM reasoning with causal structures. Extensive experiments demonstrate that CRFT substantially enhances both reasoning accuracy and interpretability across multiple model backbones. By unifying dataset curation, causal evaluation, and causality-informed learning, CausalPhys establishes a strong foundation for advancing modern VLMs toward causally grounded physical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。