构建首个融合图文的3D因果推理评测基准,助力AI理解视觉数据深层逻辑。
CAUSAL3D: A Comprehensive Benchmark for Causal Learning from Visual Data
- 设计19个3D场景数据集,融合表格与图像评估因果推理能力
- 复杂因果结构下模型性能显著下降,暴露现有方法瓶颈
- 适合研究因果推理、可信AI的学者和开发者使用
真正智能依赖于发现并利用隐藏的因果关系。尽管人工智能与计算机视觉取得进展,但缺乏评估模型从复杂视觉数据中推断潜在因果关系能力的基准。本文提出 extsc{Causal3D},一个新颖且全面的基准,将结构化数据(表格)与对应视觉表示(图像)结合,用于评估因果推理能力。该基准基于系统化框架,包含19个3D场景数据集,涵盖多样的因果关系、视角与背景,支持在不同复杂度场景下的评估。我们测试了多种先进方法,包括经典因果发现、因果表征学习及大模型(如LLMs/VLMs)。实验表明,当因果结构更复杂且无先验知识时,性能显著下降,凸显了先进方法在复杂场景中的挑战。Causal3D为推动计算机视觉中的因果推理发展提供了关键资源,并有助于在关键领域实现可信AI。
原文摘要 · Abstract (English)
True intelligence hinges on the ability to uncover and leverage hidden causal relations. Despite significant progress in AI and computer vision (CV), there remains a lack of benchmarks for assessing models' abilities to infer latent causality from complex visual data. In this paper, we introduce \textsc{\textbf{Causal3D}}, a novel and comprehensive benchmark that integrates structured data (tables) with corresponding visual representations (images) to evaluate causal reasoning. Designed within a systematic framework, Causal3D comprises 19 3D-scene datasets capturing diverse causal relations, views, and backgrounds, enabling evaluations across scenes of varying complexity. We assess multiple state-of-the-art methods, including classical causal discovery, causal representation learning, and large/vision-language models (LLMs/VLMs). Our experiments show that as causal structures grow more complex without prior knowledge, performance declines significantly, highlighting the challenges even advanced methods face in complex causal scenarios. Causal3D serves as a vital resource for advancing causal reasoning in CV and fostering trustworthy AI in critical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。