arXiv:2507.02705cs.CV2025-07NeurIPS被引 17

无需特征对齐,实现3D重建与理解的联合优化

SIU3R: Simultaneous Scene Understanding and 3D Reconstruction Beyond Feature Alignment

  • 用像素对齐的3D表示打通重建与理解任务
  • 统一语义分割等任务为可学习查询,提升3D理解能力
  • 轻量模块促进双任务协同,适合多模态智能体开发

同步场景理解与3D重建对构建端到端具身智能系统至关重要。现有方法依赖2D到3D特征对齐,导致3D理解能力受限且易丢失语义信息。为此,我们提出SIU3R——首个无需对齐的通用框架,可从无姿态图像中实现可泛化的同步理解与3D重建。该框架通过像素对齐的3D表示连接重建与理解任务,并将多种理解(分割)任务统一为一组可学习查询,实现原生3D理解,无需与2D模型对齐。为进一步促进两任务共享表示下的协作,我们深入分析其相互增益,并设计两个轻量级模块以增强交互。大量实验表明,该方法在3D重建、理解及联合任务上均达到最先进性能,验证了无对齐框架的优势与互惠设计的有效性。

原文摘要 · Abstract (English)

Simultaneous understanding and 3D reconstruction plays an important role in developing end-to-end embodied intelligent systems. To achieve this, recent approaches resort to 2D-to-3D feature alignment paradigm, which leads to limited 3D understanding capability and potential semantic information loss. In light of this, we propose SIU3R, the first alignment-free framework for generalizable simultaneous understanding and 3D reconstruction from unposed images. Specifically, SIU3R bridges reconstruction and understanding tasks via pixel-aligned 3D representation, and unifies multiple understanding (segmentation) tasks into a set of unified learnable queries, enabling native 3D understanding without the need of alignment with 2D models. To encourage collaboration between the two tasks with shared representation, we further conduct in-depth analyses of their mutual benefits, and propose two lightweight modules to facilitate their interaction. Extensive experiments demonstrate that our method achieves state-of-the-art performance not only on the individual tasks of 3D reconstruction and understanding, but also on the task of simultaneous understanding and 3D reconstruction, highlighting the advantages of our alignment-free framework and the effectiveness of the mutual benefit designs. Project page: https://insomniaaac.github.io/siu3r/

3D重建场景理解无对齐具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。