arXiv:2411.12635cs.CV2024-11被引 3

用双流结构和深度引导提升单图3D重建精度

M3D: Dual-Stream Selective State Spaces and Depth-Driven Framework for High-Fidelity Single-View 3D Reconstruction

  • 双流设计结合选择性状态空间,平衡全局与局部特征提取
  • 融合多尺度特征与深度信息,显著提升几何一致性和细节保真度
  • 适合需要高精度3D重建的虚拟现实与自动驾驶场景

从复杂场景的单张RGB图像精确重建3D物体是虚拟现实、自动驾驶和机器人领域的重要挑战。现有神经隐式3D表示方法在平衡全局与局部特征提取方面存在困难,尤其在多样化复杂环境中,导致重建精度和质量不足。本文提出M3D,一种新型单视图3D重建框架。该框架采用基于选择性状态空间的双流特征提取策略,有效平衡全局与局部特征的获取,提升场景理解与表示精度。此外,平行分支提取深度信息,实现视觉与几何特征的有效融合,增强重建质量并保留精细细节。实验表明,通过双分支特征提取融合多尺度特征与深度信息,显著提升了几何一致性与保真度,达到当前最优重建性能。

原文摘要 · Abstract (English)

The precise reconstruction of 3D objects from a single RGB image in complex scenes presents a critical challenge in virtual reality, autonomous driving, and robotics. Existing neural implicit 3D representation methods face significant difficulties in balancing the extraction of global and local features, particularly in diverse and complex environments, leading to insufficient reconstruction precision and quality. We propose M3D, a novel single-view 3D reconstruction framework, to tackle these challenges. This framework adopts a dual-stream feature extraction strategy based on Selective State Spaces to effectively balance the extraction of global and local features, thereby improving scene comprehension and representation precision. Additionally, a parallel branch extracts depth information, effectively integrating visual and geometric features to enhance reconstruction quality and preserve intricate details. Experimental results indicate that the fusion of multi-scale features with depth information via the dual-branch feature extraction significantly boosts geometric consistency and fidelity, achieving state-of-the-art reconstruction performance.

3D重建深度引导双流网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。