让3D高斯点云理解隐含意图和空间逻辑,提升智能体交互能力。
CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting

- 融合视觉语言模型与3D场景图,分离结构感知与逻辑推理。
- 在新构建的两个基准上显著超越现有方法,尤其在反事实推理上提升37%。
- 适合研究具身智能、3D场景理解与多模态推理的学者使用。
尽管3D高斯点云(3DGS)推动了开放词汇场景理解的发展,现有方法仍局限于显式查询。它们难以解析隐含意图、复杂空间约束及常识推理,无法满足实际具身交互需求。为此,我们提出3D高斯分割中的推理任务,并构建两个基准:Causal-LERF与Causal-ScanNet,系统评估常识、空间、可用性与反事实推理能力。评估显示,当前最先进方法在这些推理挑战上表现不佳。因此,我们提出CausalSplat框架,通过将视觉语言模型与3D场景图结合,解耦显式结构感知与隐式逻辑推理。大量实验表明,CausalSplat在我们的推理基准上达到领先性能,并在标准指代与开放词汇3D分割任务中展现强泛化能力。
原文摘要 · Abstract (English)
While 3D Gaussian Splatting (3DGS) has advanced open vocabulary scene understanding, existing methods remain confined to explicit queries. They struggle to interpret implicit intents, complex spatial constraints, and commonsense reasoning required for practical embodied interactions. To address this gap, we introduce the task of reasoning 3D Gaussian segmentation and construct two benchmarks, Causal-LERF and Causal-ScanNet. These benchmarks systematically evaluate commonsense, spatial, affordance, and counterfactual reasoning. Evaluations reveal that current state of the art methods perform poorly on these reasoning challenges. Therefore, we propose CausalSplat, a framework that integrates vision-language models with 3D scene graphs to disentangle explicit structural perception from implicit logical inference. Extensive experiments demonstrate that CausalSplat achieves state of the art performance on our reasoning benchmarks while showing strong generalizability on standard referring and open vocabulary 3D segmentation tasks. Project Page: https://jiayuding031020.github.io/CausalSplat
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。