无需训练,用多视角信息实现3D问答,让机器人看得懂空间关系。
ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA

- 分步拆解3D问答:选图、找物、编码空间、推理作答
- 在SQA3D上对空间问题提升明显,整体准确率50.8%
- 不需微调,适合快速部署到真实机器人场景
大型语言模型(LLMs)和视觉语言模型(VLMs)的进展为3D问答(3D-QA)带来了新可能,这是具身智能与机器人感知的关键能力。然而,现有方法大多依赖于昂贵标注的3D特化训练或微调,限制了可扩展性与实际应用。我们提出ViewMind3D,一个完全无需训练、模块化的框架,仅通过多视角图像进行3D空间推理,无需完整3D重建。该框架将3D-QA任务分解为四个可解释模块:(1) 基于问题的多视角选择,(2) 语言引导的视觉定位,(3) 鸟瞰图视角指示的空间上下文编码,(4) 基于角色的结构化答案生成。这一设计使推理过程结构清晰、鲁棒且可解释,无需模型调优。在ScanQA和SQA3D上的实验表明,ViewMind3D在训练自由与微调模型中表现相当。尤其在空间定位类问题(如SQA3D中的'What'类问题)上显著提升,整体准确率达50.8%,在ScanQA上获得73.4的CIDEr得分。结果表明,通过通用大模型的模块化协同,即可实现高效可靠的3D推理,适用于真实环境中的机器人感知。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled new possibilities for 3D question answering (3D-QA), a key capability for embodied AI and robotic perception. However, most existing methods rely on 3D-specific training or fine-tuning with costly annotations, limiting their scalability and real-world applicability. We present \textbf{ViewMind3D}, a fully training-free and modular framework for 3D spatial reasoning over multi-view observations of a scene without requiring complete 3D reconstruction. The framework decomposes the 3D-QA task into four interpretable components: (1) question-driven multi-view selection, (2) guided visual grounding with language-conditioned object cues, (3) spatial context encoding via a bird's-eye-view (BEV) viewpoint indicator, and (4) structured answer generation through role-based reasoning. This design enables structured, robust, and interpretable reasoning without requiring model tuning. Experimental results on ScanQA and SQA3D show that ViewMind3D achieves competitive performance compared to prior training-free and fine-tuned 3D-LLMs. In particular, our method improves performance on spatially grounded question types, such as ``What'' questions in SQA3D, while maintaining strong overall accuracy (50.8\%) and achieving 73.4 CIDEr on ScanQA. These results demonstrate that effective 3D reasoning can be achieved through modular orchestration of general-purpose LLMs and VLMs for robotic perception in real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。