让3D高斯点云具备语义与物理可执行性,用于真实机器人导航
Towards Physically Executable 3D Gaussian for Embodied Navigation
- 在3D高斯点云中加入物体级语义标注和碰撞体
- 新基准上比基线提升31%的未见场景导航性能
- 适合做具身智能、视觉语言导航的研究者使用
3D高斯泼溅(3DGS)具有实时逼真渲染能力,被视为缩小仿真到现实差距的有效工具。然而,它在视觉-语言导航(VLN)任务中缺乏细粒度语义和物理可执行性。为此,我们提出SAGE-3D(语义与物理对齐的3D导航高斯环境),将3DGS升级为可执行、语义与物理对齐的环境。其包含两个组件:(1) 物体中心语义定位,为3DGS添加物体级别的细粒度标注;(2) 物理感知执行连接,将碰撞物体嵌入3DGS并构建丰富的物理交互接口。我们发布了InteriorGS数据集,包含1000个带物体标注的3DGS室内场景,并引入SAGE-Bench,首个基于3DGS的VLN基准,含200万条VLN数据。实验表明,3DGS场景数据更难收敛,但泛化能力强,在VLN-CE未见任务上使基线性能提升31%。数据与代码已开源:https://sage-3d.github.io。
原文摘要 · Abstract (English)
3D Gaussian Splatting (3DGS), a 3D representation method with photorealistic real-time rendering capabilities, is regarded as an effective tool for narrowing the sim-to-real gap. However, it lacks fine-grained semantics and physical executability for Visual-Language Navigation (VLN). To address this, we propose SAGE-3D (Semantically and Physically Aligned Gaussian Environments for 3D Navigation), a new paradigm that upgrades 3DGS into an executable, semantically and physically aligned environment. It comprises two components: (1) Object-Centric Semantic Grounding, which adds object-level fine-grained annotations to 3DGS; and (2) Physics-Aware Execution Jointing, which embeds collision objects into 3DGS and constructs rich physical interfaces. We release InteriorGS, containing 1K object-annotated 3DGS indoor scene data, and introduce SAGE-Bench, the first 3DGS-based VLN benchmark with 2M VLN data. Experiments show that 3DGS scene data is more difficult to converge, while exhibiting strong generalizability, improving baseline performance by 31% on the VLN-CE Unseen task. Our data and code are available at: https://sage-3d.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。