用动态高斯溅射提升仿真真实感,让智能体学会与真人互动。
Habitat-GS: A High-Fidelity Navigation Simulator with Dynamic Gaussian Splatting

- 用高斯溅射技术实现实时逼真渲染,支持复杂场景导入
- 引入高斯人像模块,兼具视觉真实与导航障碍功能
- 训练出的智能体在跨域任务中表现更优,适合真实场景部署
构建基于3D高斯溅射(3DGS)的导航模拟器Habitat-GS,扩展自Habitat-Sim,实现高保真视觉渲染与动态人类建模。系统采用实时3DGS渲染器,支持从多种来源导入可扩展的3DGS资产;提出高斯人像模块,使每个虚拟人物既具备照片级视觉效果,又能作为有效导航障碍物,促进智能体学习人机交互行为。点目标导航实验表明,基于3DGS场景训练的智能体具有更强跨域泛化能力,混合域训练策略最优。人像感知导航评估验证了高斯人像在真实环境适应性上的有效性。性能基准测试显示系统在不同场景复杂度和人像数量下均具可扩展性。
原文摘要 · Abstract (English)
Training embodied AI agents depends critically on the visual fidelity of simulation environments and the ability to model dynamic humans. Current simulators rely on mesh-based rasterization with limited visual realism, and their support for dynamic human avatars, where available, is constrained to mesh representations, hindering agent generalization to human-populated real-world scenarios. We present Habitat-GS, a navigation-centric embodied AI simulator extended from Habitat-Sim that integrates 3D Gaussian Splatting scene rendering and drivable gaussian avatars while maintaining full compatibility with the Habitat ecosystem. Our system implements a 3DGS renderer for real-time photorealistic rendering and supports scalable 3DGS asset import from diverse sources. For dynamic human modeling, we introduce a gaussian avatar module that enables each avatar to simultaneously serve as a photorealistic visual entity and an effective navigation obstacle, allowing agents to learn human-aware behaviors in realistic settings. Experiments on point-goal navigation demonstrate that agents trained on 3DGS scenes achieve stronger cross-domain generalization, with mixed-domain training being the most effective strategy. Evaluations on avatar-aware navigation further confirm that gaussian avatars enable effective human-aware navigation. Finally, performance benchmarks validate the system's scalability across varying scene complexity and avatar counts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。