用2D与3D跨模态交互提升纯摄像头4D占位预测性能
InterOCF: Spatio-Temporal 2D-3D Interaction for Camera-Only 4D Occupancy Forecasting

- 设计2D-3D双分支时空交互模块,同步建模图像与体素动态
- 在nuScenes等3个数据集上平均性能超越现有方法
- 适合关注自动驾驶感知与多视角时序建模的研究者
仅依赖多视角图像的4D占位预测对自动驾驶安全至关重要。尽管现有方法表现良好,但输入帧间强时空关联仍未充分挖掘,制约了未来场景预测性能。为此,本文提出InterOCF框架,联合建模3D体素表示与多视图分割序列的时序动态,并显式实现2D与3D分支间的特征交互。核心包含三个组件:1)3D时空(3DST)模块,从历史体素状态学习体积动态以预测未来状态;2)2D时空(2DST)模块,通过辅助多视图时序分割任务增强语义动态建模;3)时空交互建模(STIM)模块,促进2D与3D表示间的信息交换。在nuScenes、Lyft-Level5及nuScenes-Occupancy数据集上的实验表明,InterOCF持续优于现有基线方法。
原文摘要 · Abstract (English)
Camera-only 4D occupancy forecasting enables autonomous vehicles to predict future 3D semantic scenes solely from historical multi-view images, which is critical for driving safety. Even though current methods have achieved good performance, the strong spatial-temporal modeling between the input multi-view frames is still underexplored, which limits the performance of those methods in future 4D forecasting. To address this gap, we introduce a novel framework, InterOCF, for 4D occupancy forecasting that jointly models temporal dynamics in both 3D voxel-based representations and multi-view segmentation sequences, while explicitly incorporating feature interaction between the 2D and 3D branches. Our framework incorporates three core components: 1) A 3D Spatio-Temporal (3DST) module that learns volumetric dynamics from historical voxel states to predict future voxel states; 2) A 2D Spatio-Temporal (2DST) module employing an auxiliary multi-view temporal segmentation forecasting task to enhance temporal semantic dynamics; 3) A Spatio-Temporal Interaction Modeling (STIM) module that enables feature interaction between 2D and 3D representations. Experiments on the nuScenes, Lyft-Level5, and nuScenes-Occupancy datasets show that InterOCF consistently outperforms existing baseline approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。