用3D高斯点云融合语义与几何,实现大场景实时全景重建。
Ψ-Map: Panoptic Surface Integrated Mapping Enables Real2Sim Transfer
- 用激光数据约束高斯混合模型,结合2D高斯面元提升表面精度。
- 端到端架构避免逐帧分割误差累积,实现全局一致的语义理解。
- 优化渲染流程,实现在大场景中超过40帧/秒的实时推理。
开放词汇全景重建对先进机器人感知与仿真至关重要。然而,现有基于3D高斯点云(3DGS)的方法在大规模场景中难以同时实现几何精度、连贯的全景理解与实时推理。本文提出一种集成几何增强、端到端全景学习与高效渲染的综合框架。首先,利用激光雷达数据构建平面约束的多模态高斯混合模型(GMM),以2D高斯面元作为地图表示,实现高精度表面对齐与连续几何监督。其次,为克服传统多阶段全景分割中误差累积与繁琐跨帧关联的问题,设计基于查询引导的端到端学习架构,通过视锥内局部交叉注意力机制,将2D掩码特征直接提升至3D空间,实现全局一致的全景理解。最后,针对高维语义特征带来的计算瓶颈,引入精确瓦片交集与Top-K硬选择策略优化渲染管道。实验表明,该系统在大规模场景中实现了更优的几何与全景重建质量,同时保持超过40 FPS的推理速率,满足机器人控制回路的实时性要求。
原文摘要 · Abstract (English)
Open-vocabulary panoptic reconstruction is essential for advanced robotics perception and simulation. However, existing methods based on 3D Gaussian Splatting (3DGS) often struggle to simultaneously achieve geometric accuracy, coherent panoptic understanding, and real-time inference frequency in large-scale scenes. In this paper, we propose a comprehensive framework that integrates geometric reinforcement, end-to-end panoptic learning, and efficient rendering. First, to ensure physical realism in large-scale environments, we leverage LiDAR data to construct plane-constrained multimodal Gaussian Mixture Models (GMMs) and employ 2D Gaussian surfels as the map representation, enabling high-precision surface alignment and continuous geometric supervision. Building upon this, to overcome the error accumulation and cumbersome cross-frame association inherent in traditional multi-stage panoptic segmentation pipelines, we design a query-guided end-to-end learning architecture. By utilizing a local cross-attention mechanism within the view frustum, the system lifts 2D mask features directly into 3D space, achieving globally consistent panoptic understanding. Finally, addressing the computational bottlenecks caused by high-dimensional semantic features, we introduce Precise Tile Intersection and a Top-K Hard Selection strategy to optimize the rendering pipeline. Experimental results demonstrate that our system achieves superior geometric and panoptic reconstruction quality in large-scale scenes while maintaining an inference rate exceeding 40 FPS, meeting the real-time requirements of robotic control loops.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。