arXiv:2510.16714cs.CVcs.AI2025-10中稿 · ICLR被引 7

让3D场景理解像人一样一步步推理,提升问答准确性。

SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D Scenes

  • 将复杂推理拆解为可管理步骤,结合视觉线索逐步推导
  • 在18.5万条高质量数据上训练,显著提升问答一致性
  • 适合做3D场景理解、智能机器人导航的研究者

现有3D大语言模型在实现有依据的问答方面仍存在困难,主要因对人类式场景-物体关联推理机制研究不足。本文提出一种新型框架,引入3D场景中的有根据思维链方法(SCENECOT),将复杂任务分解为简单子问题,并基于多模态专家模块构建相应视觉线索。为支持该方法,我们构建了首个大规模有根据思维链数据集SCENECOT-185K,包含18.5万条高质量样本。在多个复杂3D场景推理基准上的实验表明,新框架在保持高问答一致性的前提下取得优异表现。据我们所知,这是首次成功将思维链推理应用于3D场景理解,实现了类似人类的分步推理,具备向更广泛3D场景理解场景扩展的潜力。

原文摘要 · Abstract (English)

Existing research on 3D Large Language Models (LLMs) still struggles to achieve grounded question-answering, primarily due to the under-exploration of the mechanism of human-like scene-object grounded reasoning. This paper bridges the gap by presenting a novel framework. We first introduce a grounded Chain-of-Thought reasoning method in 3D scenes (SCENECOT), decoupling a complex reasoning task into simpler and manageable problems, and building corresponding visual clues based on multimodal expert modules. To enable such a method, we develop SCENECOT-185K, the first large-scale grounded CoT reasoning dataset, consisting of 185K high-quality instances. Extensive experiments across various complex 3D scene reasoning benchmarks demonstrate that our new framework achieves strong performance with high grounding-QA coherence. To the best of our knowledge, this is the first successful application of CoT reasoning to 3D scene understanding, enabling step-by-step human-like reasoning and showing potential for extension to broader 3D scene understanding scenarios.

3D理解思维链多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。