通过因果损失实现端到端训练,提升3D语义占据预测的准确性与鲁棒性。
Semantic Causality-Aware Vision-Based 3D Occupancy Prediction

- 设计因果损失,使2D到3D转换全流程可微分
- 在Occ3D上达到当前最优性能,对相机扰动更鲁棒
- 适合自动驾驶、机器人等需要精准3D感知的场景
基于视觉的3D语义占据预测是3D视觉中的关键任务,融合了体素化3D重建与语义理解。现有方法通常采用模块化流水线,各模块独立优化或使用预设输入,导致误差累积。本文提出一种新型因果损失,实现2D到3D转换流水线的全局端到端监督。该损失基于2D到3D语义因果性,调控从3D体素表示向2D特征的梯度回传,使整个流程可微分,统一学习过程,使以往不可训练的组件完全可学习。在此基础上,提出语义因果感知的2D到3D转换框架,包含三个由因果损失引导的组件:通道分组提升用于自适应语义映射,可学习相机偏移增强对相机扰动的鲁棒性,归一化卷积实现高效特征传播。大量实验表明,该方法在Occ3D基准上达到领先性能,显著提升对相机扰动的鲁棒性与2D到3D语义一致性。
原文摘要 · Abstract (English)
Vision-based 3D semantic occupancy prediction is a critical task in 3D vision that integrates volumetric 3D reconstruction with semantic understanding. Existing methods, however, often rely on modular pipelines. These modules are typically optimized independently or use pre-configured inputs, leading to cascading errors. In this paper, we address this limitation by designing a novel causal loss that enables holistic, end-to-end supervision of the modular 2D-to-3D transformation pipeline. Grounded in the principle of 2D-to-3D semantic causality, this loss regulates the gradient flow from 3D voxel representations back to the 2D features. Consequently, it renders the entire pipeline differentiable, unifying the learning process and making previously non-trainable components fully learnable. Building on this principle, we propose the Semantic Causality-Aware 2D-to-3D Transformation, which comprises three components guided by our causal loss: Channel-Grouped Lifting for adaptive semantic mapping, Learnable Camera Offsets for enhanced robustness against camera perturbations, and Normalized Convolution for effective feature propagation. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the Occ3D benchmark, demonstrating significant robustness to camera perturbations and improved 2D-to-3D semantic consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。