提出新方法提升3D语义占据与运动预测精度和效率
ALOcc: Adaptive Lifting-Based 3D Semantic Occupancy and Cost Volume-Based Flow Predictions

- 自适应提升机制结合去噪,增强2D到3D特征转换鲁棒性
- 联合优化原型,解决长尾类别问题,提升语义一致性
- 基于BEV的代价体积设计,支持高效多任务联合预测
3D语义占据与运动预测是时空场景理解的基础。本文提出一种纯视觉框架,实现三项改进:首先,引入考虑遮挡的自适应提升机制并融合深度去噪,增强2D到3D特征转换的鲁棒性,降低对深度先验的依赖;其次,通过联合优化原型并采用置信度与类别感知采样,强化3D-2D语义一致性,缓解长尾类别问题;第三,设计以鸟瞰图(BEV)为中心的代价体积,显式关联语义与运动特征,并采用混合分类-回归监督策略,适配不同尺度的运动变化。所提纯卷积架构在多个基准上实现新SOTA性能,涵盖语义占据及联合占据-语义-运动预测。同时提供一系列效率-性能权衡模型,其实时版本在速度与准确率上均超越现有实时方法,具备实际应用价值。
原文摘要 · Abstract (English)
3D semantic occupancy and flow prediction are fundamental to spatiotemporal scene understanding. This paper proposes a vision-based framework with three targeted improvements. First, we introduce an occlusion-aware adaptive lifting mechanism incorporating depth denoising. This enhances the robustness of 2D-to-3D feature transformation while mitigating reliance on depth priors. Second, we enforce 3D-2D semantic consistency via jointly optimized prototypes, using confidence- and category-aware sampling to address the long-tail classes problem. Third, to streamline joint prediction, we devise a BEV-centric cost volume to explicitly correlate semantic and flow features, supervised by a hybrid classification-regression scheme that handles diverse motion scales. Our purely convolutional architecture establishes new SOTA performance on multiple benchmarks for both semantic occupancy and joint occupancy semantic-flow prediction. We also present a family of models offering a spectrum of efficiency-performance trade-offs. Our real-time version exceeds all existing real-time methods in speed and accuracy, ensuring its practical viability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。