arXiv:2503.04852cs.CVcs.LG2025-03被引 7

构建首个融合图文的3D因果推理评测基准,助力AI理解视觉数据深层逻辑。

CAUSAL3D: A Comprehensive Benchmark for Causal Learning from Visual Data

  • 设计19个3D场景数据集,融合表格与图像评估因果推理能力
  • 复杂因果结构下模型性能显著下降,暴露现有方法瓶颈
  • 适合研究因果推理、可信AI的学者和开发者使用

真正智能依赖于发现并利用隐藏的因果关系。尽管人工智能与计算机视觉取得进展,但缺乏评估模型从复杂视觉数据中推断潜在因果关系能力的基准。本文提出 extsc{Causal3D},一个新颖且全面的基准,将结构化数据(表格)与对应视觉表示(图像)结合,用于评估因果推理能力。该基准基于系统化框架,包含19个3D场景数据集,涵盖多样的因果关系、视角与背景,支持在不同复杂度场景下的评估。我们测试了多种先进方法,包括经典因果发现、因果表征学习及大模型(如LLMs/VLMs)。实验表明,当因果结构更复杂且无先验知识时,性能显著下降,凸显了先进方法在复杂场景中的挑战。Causal3D为推动计算机视觉中的因果推理发展提供了关键资源,并有助于在关键领域实现可信AI。

原文摘要 · Abstract (English)

True intelligence hinges on the ability to uncover and leverage hidden causal relations. Despite significant progress in AI and computer vision (CV), there remains a lack of benchmarks for assessing models' abilities to infer latent causality from complex visual data. In this paper, we introduce \textsc{\textbf{Causal3D}}, a novel and comprehensive benchmark that integrates structured data (tables) with corresponding visual representations (images) to evaluate causal reasoning. Designed within a systematic framework, Causal3D comprises 19 3D-scene datasets capturing diverse causal relations, views, and backgrounds, enabling evaluations across scenes of varying complexity. We assess multiple state-of-the-art methods, including classical causal discovery, causal representation learning, and large/vision-language models (LLMs/VLMs). Our experiments show that as causal structures grow more complex without prior knowledge, performance declines significantly, highlighting the challenges even advanced methods face in complex causal scenarios. Causal3D serves as a vital resource for advancing causal reasoning in CV and fostering trustworthy AI in critical domains.

因果推理3D场景评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。