用稀疏感知量化压缩3D语义占据,通信量降82倍仍保精度。
Sparse-Aware Vector Quantization for Bandwidth-Efficient Collaborative 3D Semantic Occupancy Prediction

- 基于3D场景稀疏性设计量化机制,只传关键区域信息
- 通信量减少82倍,3D几何结构完整保留
- 适合自动驾驶多车协同感知,低带宽场景部署
协同感知通过多车交换互补感知信息扩展单智能体感知能力,但存在感知增益与通信开销的固有权衡,尤其在依赖精细空间结构的3D语义占据预测中更为显著。现有方法通常将3D特征压缩为2D,导致严重空间信息损失,或传输密集3D表示,阻碍实际部署。为此,我们提出一种高效通信的协同向量量化语义占据预测框架(VQSOP)。VQSOP采用稀疏感知向量量化(SAVQ)机制,利用3D场景稀疏性对关键区域进行紧凑编码,大幅降低通信开销并保持完整几何上下文。此外,为增强结构一致性和特征连续性,设计双分支自适应空间精修(ASR)模块,动态融合局部高频细节与全局语义信息。大量实验表明,该方法在通信量减少82倍的同时达到领先性能。
原文摘要 · Abstract (English)
Collaborative perception extends single-agent perception by enabling multiple vehicles to exchange complementary perceptual information. However, it introduces an inherent trade-off between perception gain and communication overhead, which is particularly severe for 3D semantic occupancy prediction that relies on fine-grained spatial structures. Existing methods typically compress 3D features into 2D, causing severe spatial information loss, or transmit dense 3D representations, hindering real-world deployment. To overcome these limitations, we propose a bandwidth-efficient collaborative Vector Quantization Semantic Occupancy Prediction (VQSOP) framework. VQSOP employs a Sparse-Aware Vector Quantization (SAVQ) mechanism that exploits 3D scene sparsity to compactly encode informative regions, drastically reducing communication overhead while preserving complete geometric context. Furthermore, to enhance structural consistency and feature continuity, we design a Dual-Branch Adaptive Spatial Refinement (ASR) module that dynamically fuses local high-frequency details with broad contextual semantics. Extensive experiments demonstrate that our approach achieves state-of-the-art performance while reducing communication volume by up to 82x.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。