arXiv:2503.09941cs.CVcs.AI2025-03被引 1

用3D高斯和稀疏点融合提升3D语义占用预测精度

TGP: Two-modal occupancy prediction with 3D Gaussian and sparse points for 3D Environment Awareness

  • 结合3D高斯与稀疏点的双模态建模方法
  • 在Occ3DnuScenes上实现更高的IoU指标
  • 适合需要精确环境感知的机器人与自动驾驶场景

3D语义占用感知因能提供更真实的几何信息并更好融入下游任务,已成为机器人与自动驾驶环境感知的研究热点。通过预测环境中3D空间的占用情况,可有效提升场景理解的能力与鲁棒性。然而,现有方法多基于体素或点云:体素化导致空间信息丢失,点云虽保留位置信息却难以表达体积结构。为此,本文提出一种基于3D高斯集合与稀疏点的双模态预测方法,兼顾空间位置与体积结构信息,显著提升语义占用预测精度。具体地,采用Transformer架构,以3D高斯集合、稀疏点及查询作为输入,通过多层Transformer结构,增强查询与高斯集合协同参与预测,并设计自适应融合机制整合双模态语义输出,生成最终结果。此外,每层动态优化点云,提升定位精度。在Occ3DnuScenes数据集上的实验表明,该方法在基于IoU的指标上表现优异。

原文摘要 · Abstract (English)

3D semantic occupancy has rapidly become a research focus in the fields of robotics and autonomous driving environment perception due to its ability to provide more realistic geometric perception and its closer integration with downstream tasks. By performing occupancy prediction of the 3D space in the environment, the ability and robustness of scene understanding can be effectively improved. However, existing occupancy prediction tasks are primarily modeled using voxel or point cloud-based approaches: voxel-based network structures often suffer from the loss of spatial information due to the voxelization process, while point cloud-based methods, although better at retaining spatial location information, face limitations in representing volumetric structural details. To address this issue, we propose a dual-modal prediction method based on 3D Gaussian sets and sparse points, which balances both spatial location and volumetric structural information, achieving higher accuracy in semantic occupancy prediction. Specifically, our method adopts a Transformer-based architecture, taking 3D Gaussian sets, sparse points, and queries as inputs. Through the multi-layer structure of the Transformer, the enhanced queries and 3D Gaussian sets jointly contribute to the semantic occupancy prediction, and an adaptive fusion mechanism integrates the semantic outputs of both modalities to generate the final prediction results. Additionally, to further improve accuracy, we dynamically refine the point cloud at each layer, allowing for more precise location information during occupancy prediction. We conducted experiments on the Occ3DnuScenes dataset, and the experimental results demonstrate superior performance of the proposed method on IoU based metrics.

3D感知语义占用高斯过程自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。