让AI通过重建几何结构理解3D场景,效果显著提升。
Reg3D: Reconstructive Geometry Instruction Tuning for 3D Scene Understanding
- 用重建几何结构代替单纯描述,引导模型学习空间关系。
- 在4个数据集上性能大幅领先,验证了新训练范式有效性。
- 适合研究3D视觉理解、多模态模型的开发者与研究人员。
大型多模态模型(LMMs)在2D视觉理解方面取得了显著进展,但将其能力扩展到3D场景理解仍面临重大挑战。现有方法主要依赖纯文本监督,无法提供学习稳健3D空间表征所需的几何约束。本文提出Reg3D,一种重构式几何指令微调框架,通过将几何感知监督直接融入训练过程来解决此问题。核心思想是:有效的3D理解需要重建底层几何结构,而非仅进行描述。不同于仅在输入层注入3D信息的方法,Reg3D采用双监督范式,将3D几何信息同时作为输入和显式学习目标。我们设计了互补的对象级与帧级重建任务,结合双编码器架构,强制几何一致性以促进空间推理能力的发展。在ScanQA、Scan2Cap、ScanRefer和SQA3D上的大量实验表明,Reg3D实现了显著的性能提升,确立了新型空间感知多模态模型的训练范式。
原文摘要 · Abstract (English)
The rapid development of Large Multimodal Models (LMMs) has led to remarkable progress in 2D visual understanding; however, extending these capabilities to 3D scene understanding remains a significant challenge. Existing approaches predominantly rely on text-only supervision, which fails to provide the geometric constraints required for learning robust 3D spatial representations. In this paper, we introduce Reg3D, a novel Reconstructive Geometry Instruction Tuning framework that addresses this limitation by incorporating geometry-aware supervision directly into the training process. Our key insight is that effective 3D understanding necessitates reconstructing underlying geometric structures rather than merely describing them. Unlike existing methods that inject 3D information solely at the input level, Reg3D adopts a dual-supervision paradigm that leverages 3D geometric information both as input and as explicit learning targets. Specifically, we design complementary object-level and frame-level reconstruction tasks within a dual-encoder architecture, enforcing geometric consistency to encourage the development of spatial reasoning capabilities. Extensive experiments on ScanQA, Scan2Cap, ScanRefer, and SQA3D demonstrate that Reg3D delivers substantial performance improvements, establishing a new training paradigm for spatially aware multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。