SAB3R统一3D重建与开放词汇语义分割,实现单次前向生成带语义的点云
SAB3R: Semantic-Augmented Backbone in 3D Reconstruction
- 用轻量级蒸馏将2D视觉模型语义特征注入3D重建框架
- 在Map and Locate任务上优于分离部署的MASt3R+CLIP组合
- 适合构建具语义理解能力的现实世界智能体系统
我们提出新任务Map and Locate,融合开放词汇分割与3D重建:从无姿态视频生成点云,并基于自然语言查询分割物体实例。该任务是实现真实场景具身智能的关键步骤。为此,我们提出SAB3R基线模型,基于MASt3R改进,通过轻量级蒸馏策略,将CLIP和DINOv2等2D视觉骨干网络的密集像素级语义特征迁移至3D重建流程。无需额外冻结网络,模型可在一次前向传播中生成像素级语义特征并构建连贯点图。相比独立使用MASt3R与CLIP,SAB3R在Map and Locate基准测试中表现更优。此外,我们在2D语义分割与3D任务上全面验证了其有效性。
原文摘要 · Abstract (English)
We introduce a new task, Map and Locate, which unifies the traditionally distinct objectives of open-vocabulary segmentation - detecting and segmenting object instances based on natural language queries - and 3D reconstruction, the process of estimating a scene's 3D structure from visual inputs. Specifically, Map and Locate involves generating a point cloud from an unposed video and segmenting object instances based on open-vocabulary queries. This task serves as a critical step toward real-world embodied AI applications and introduces a practical task that bridges reconstruction, recognition and reorganization. To tackle this task, we introduce a simple yet effective baseline, which we denote as SAB3R. Our approach builds upon MASt3R, a recent breakthrough in 3D computer vision, and incorporates a lightweight distillation strategy. This method transfers dense, per-pixel semantic features from 2D vision backbones (eg, CLIP and DINOv2) to enhance MASt3R's capabilities. Without introducing any auxiliary frozen networks, our model generates per-pixel semantic features and constructs cohesive point maps in a single forward pass. Compared to separately deploying MASt3R and CLIP, our unified model, SAB3R, achieves superior performance on the Map and Locate benchmark. Furthermore, we evaluate SAB3R on both 2D semantic segmentation and 3D tasks to comprehensively validate its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。