用语言辅助对齐3D场景图,提升机器人导航的准确性
SGAligner++: Cross-Modal Language-Aided 3D Scene Graph Alignment
- 融合语言与点云数据,构建统一嵌入空间实现跨模态对齐
- 在低重叠和噪声环境下仍保持40%性能优势
- 适合需要高鲁棒性定位与重建的机器人应用
3D场景图对齐是机器人导航与具身感知的重要前置步骤。现有方法多依赖单一模态点云数据,在输入不完整或含噪声时表现不佳。本文提出SGAligner++,一种基于语言辅助的跨模态3D场景图对齐框架。通过学习统一联合嵌入空间,该方法能够在部分重叠及异构模态间实现精准对齐,即使在传感器噪声和低重叠条件下依然有效。采用轻量级单模态编码器与注意力融合机制,显著提升视觉定位、3D重建与导航等任务的场景理解能力,同时保持高可扩展性与低计算开销。在真实世界数据集上的大量实验表明,SGAligner++在噪声重建任务中相比现有最优方法性能提升最高达40%,并具备良好的跨模态泛化能力。
原文摘要 · Abstract (English)
Aligning 3D scene graphs is a crucial initial step for several applications in robot navigation and embodied perception. Current methods in 3D scene graph alignment often rely on single-modality point cloud data and struggle with incomplete or noisy input. We introduce SGAligner++, a cross-modal, language-aided framework for 3D scene graph alignment. Our method addresses the challenge of aligning partially overlapping scene observations across heterogeneous modalities by learning a unified joint embedding space, enabling accurate alignment even under low-overlap conditions and sensor noise. By employing lightweight unimodal encoders and attention-based fusion, SGAligner++ enhances scene understanding for tasks such as visual localization, 3D reconstruction, and navigation, while ensuring scalability and minimal computational overhead. Extensive evaluations on real-world datasets demonstrate that SGAligner++ outperforms state-of-the-art methods by up to 40% on noisy real-world reconstructions, while enabling cross-modal generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。