arXiv:2512.03422cs.ROcs.CV2025-12被引 43

对比3D场景表征方法,探索机器人最佳选择。

What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models

  • 按感知、建图、定位等模块分类对比点云、SDF、NeRF等方法
  • 神经渲染与基础模型更适配语义理解与智能决策
  • 适合研究3D机器人感知与未来系统设计的学者

本文全面综述了机器人领域现有的3D场景表征方法,涵盖传统表示如点云、体素、有符号距离函数(SDF)和场景图,以及新兴的神经表示如神经辐射场(NeRF)、3D高斯溅射(3DGS)和基础模型。尽管当前SLAM与定位系统主要依赖稀疏表示如点云和体素,但密集场景表示在导航与避障等下游任务中具有关键作用。神经表示如NeRF、3DGS和基础模型能有效融合高层语义特征与语言先验,支持更全面的3D场景理解与具身智能。本文将机器人核心模块分为五类(感知、建图、定位、导航、操作),先介绍各类表征的标准形式,并比较其在不同模块中的优劣。围绕‘机器人最佳3D场景表征是什么’这一核心问题展开讨论,展望3D基础模型如何成为未来机器人的统一解决方案,并探讨其实现中的挑战。文章旨在为新老研究者提供有价值的参考,推动3D场景表征在机器人中的应用。开源项目已发布于GitHub,将持续更新新技术。

原文摘要 · Abstract (English)

In this paper, we provide a comprehensive overview of existing scene representation methods for robotics, covering traditional representations such as point clouds, voxels, signed distance functions (SDF), and scene graphs, as well as more recent neural representations like Neural Radiance Fields (NeRF), 3D Gaussian Splatting (3DGS), and the emerging Foundation Models. While current SLAM and localization systems predominantly rely on sparse representations like point clouds and voxels, dense scene representations are expected to play a critical role in downstream tasks such as navigation and obstacle avoidance. Moreover, neural representations such as NeRF, 3DGS, and foundation models are well-suited for integrating high-level semantic features and language-based priors, enabling more comprehensive 3D scene understanding and embodied intelligence. In this paper, we categorized the core modules of robotics into five parts (Perception, Mapping, Localization, Navigation, Manipulation). We start by presenting the standard formulation of different scene representation methods and comparing the advantages and disadvantages of scene representation across different modules. This survey is centered around the question: What is the best 3D scene representation for robotics? We then discuss the future development trends of 3D scene representations, with a particular focus on how the 3D Foundation Model could replace current methods as the unified solution for future robotic applications. The remaining challenges in fully realizing this model are also explored. We aim to offer a valuable resource for both newcomers and experienced researchers to explore the future of 3D scene representations and their application in robotics. We have published an open-source project on GitHub and will continue to add new works and technologies to this project.

3D场景表示机器人感知基础模型神经渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。