用原型映射提升低分辨率查询的3D占位预测精度
3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation

- 通过图像片段原型映射,增强低分辨率2D查询的3D信息
- 在Occ3D和SemanticKITTI上实现优于基线的性能
- 仅需原分辨率25%的查询量,仍保持竞争力
相机基础的3D占位预测中,体素查询分辨率显著影响视图变换质量。然而,计算约束与实时部署需求要求使用较小的查询分辨率,这不可避免导致信息丢失。因此,在有限查询尺寸下编码并保留丰富视觉细节,并确保3D占位的完整表征至关重要。为此,我们提出ProtoOcc,一种新型占位网络,利用视图变换中聚类图像片段的原型来增强低分辨率上下文。具体而言,将2D原型映射到3D体素查询上,编码高层视觉几何结构,弥补因查询分辨率降低带来的空间信息损失。此外,我们设计了多视角解码策略,高效地将密集压缩的视觉线索解耦为高维3D占位场景。在Occ3D和SemanticKITTI基准上的实验结果证明了所提方法的有效性,显著优于基线。更重要的是,ProtoOcc在体素分辨率降低75%的情况下,仍能实现与基线相当的性能。
原文摘要 · Abstract (English)
The resolution of voxel queries significantly influences the quality of view transformation in camera-based 3D occupancy prediction. However, computational constraints and the practical necessity for real-time deployment require smaller query resolutions, which inevitably leads to an information loss. Therefore, it is essential to encode and preserve rich visual details within limited query sizes while ensuring a comprehensive representation of 3D occupancy. To this end, we introduce ProtoOcc, a novel occupancy network that leverages prototypes of clustered image segments in view transformation to enhance low-resolution context. In particular, the mapping of 2D prototypes onto 3D voxel queries encodes high-level visual geometries and complements the loss of spatial information from reduced query resolutions. Additionally, we design a multi-perspective decoding strategy to efficiently disentangle the densely compressed visual cues into a high-dimensional 3D occupancy scene. Experimental results on both Occ3D and SemanticKITTI benchmarks demonstrate the effectiveness of the proposed method, showing clear improvements over the baselines. More importantly, ProtoOcc achieves competitive performance against the baselines even with 75\% reduced voxel resolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。