arXiv:2508.10936cs.CVcs.RO2025-08AAAI被引 3

用稀疏高斯点实现车联3D语义占位预测,通信量减三成仍更准。

Vision-Only Gaussian Splatting for Collaborative Semantic Occupancy Prediction

  • 共享并融合中间高斯原语,实现跨车协同去重降噪。
  • 在仅34.6%通信量下仍达mIoU提升1.9,比基线高22.41点IoU。
  • 无需深度监督,适合低带宽车路协同场景,尤其适配智能驾驶系统。

车联感知通过信息共享可克服遮挡问题,并扩展单机系统有限的感知范围。现有纯视觉3D语义占位预测方法多依赖密集3D体素,导致通信开销大;或使用2D平面特征,需精确深度估计或额外监督,限制了其在协同场景的应用。为此,我们提出首个基于稀疏3D语义高斯点阵的车联3D语义占位预测方法。通过共享与融合中间高斯原语,该方法具备三大优势:基于邻域的跨车融合可去除重复项并抑制噪声/不一致原语;每个原语联合编码几何与语义,降低对深度监督的依赖,支持简单刚性对齐;稀疏、以物体为中心的消息传递在保留结构信息的同时显著降低通信量。大量实验表明,本方法在mIoU上优于单机感知和基线协同方法8.42和3.28点,在IoU上分别提升5.11和22.41点。当进一步减少传输高斯数时,仍保持mIoU提升1.9,通信量仅为34.6%,证明其在低带宽条件下的鲁棒性能。

原文摘要 · Abstract (English)

Collaborative perception enables connected vehicles to share information, overcoming occlusions and extending the limited sensing range inherent in single-agent (non-collaborative) systems. Existing vision-only methods for 3D semantic occupancy prediction commonly rely on dense 3D voxels, which incur high communication costs, or 2D planar features, which require accurate depth estimation or additional supervision, limiting their applicability to collaborative scenarios. To address these challenges, we propose the first approach leveraging sparse 3D semantic Gaussian splatting for collaborative 3D semantic occupancy prediction. By sharing and fusing intermediate Gaussian primitives, our method provides three benefits: a neighborhood-based cross-agent fusion that removes duplicates and suppresses noisy or inconsistent Gaussians; a joint encoding of geometry and semantics in each primitive, which reduces reliance on depth supervision and allows simple rigid alignment; and sparse, object-centric messages that preserve structural information while reducing communication volume. Extensive experiments demonstrate that our approach outperforms single-agent perception and baseline collaborative methods by +8.42 and +3.28 points in mIoU, and +5.11 and +22.41 points in IoU, respectively. When further reducing the number of transmitted Gaussians, our method still achieves a +1.9 improvement in mIoU, using only 34.6% communication volume, highlighting robust performance under limited communication budgets.

3D占位车联感知高斯点阵协同推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。