arXiv:2511.07813cs.CVcs.AI2025-11AAAI被引 1

无需训练,仅用稀疏图像实现高效3D场景解析与智能推理

Sparse3DPR: Training-Free 3D Hierarchical Scene Parsing and Task-Adaptive Subgraph Reasoning from Sparse RGB Views

  • 构建分层平面增强的场景图,以平面结构为空间锚点提升推理清晰度
  • 动态提取任务相关子图,减少干扰信息,提升推理速度78.2%、准确率28.7%提升
  • 适用于开放词汇场景理解,适合实际部署且具备强泛化能力

近期,大语言模型(LLMs)被广泛探索用于3D场景理解。其中,无训练方法因其灵活性和泛化性备受关注,但通常在实际部署中面临精度与效率的挑战。为此,我们提出Sparse3DPR,一种新颖的无训练框架,支持开放式场景理解,仅需稀疏视角的RGB输入。具体地,我们引入分层平面增强的场景图,以主导平面结构作为空间锚点,支持开放词汇,使推理链更清晰、高层推断更可靠。此外,设计任务自适应子图提取方法,动态过滤与查询无关信息,降低上下文噪声,提升3D场景推理效率与准确性。实验表明,Sparse3DPR在Space3D-Bench上相较ConceptGraphs实现28.7%的EM@1提升和78.2%的速度加速;在ScanQA上表现接近训练型方法,真实世界实验也验证其鲁棒性与泛化能力。

原文摘要 · Abstract (English)

Recently, large language models (LLMs) have been explored widely for 3D scene understanding. Among them, training-free approaches are gaining attention for their flexibility and generalization over training-based methods. However, they typically struggle with accuracy and efficiency in practical deployment. To address the problems, we propose Sparse3DPR, a novel training-free framework for open-ended scene understanding, which leverages the reasoning capabilities of pre-trained LLMs and requires only sparse-view RGB inputs. Specifically, we introduce a hierarchical plane-enhanced scene graph that supports open vocabulary and adopts dominant planar structures as spatial anchors, which enables clearer reasoning chains and more reliable high-level inferences. Furthermore, we design a task-adaptive subgraph extraction method to filter query-irrelevant information dynamically, reducing contextual noise and improving 3D scene reasoning efficiency and accuracy. Experimental results demonstrate the superiority of Sparse3DPR, which achieves a 28.7% EM@1 improvement and a 78.2% speedup compared with ConceptGraphs on the Space3D-Bench. Moreover, Sparse3DPR obtains comparable performance to training-based methods on ScanQA, with additional real-world experiments confirming its robustness and generalization capability.

3D理解无训练场景图推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。