提出GAST模型,提升4D占用预测的几何保真度与时间一致性。
Geometry-Aware Spatio-Temporal Context Modeling for 4D Occupancy Forecasting

- 分阶段显式-隐式生成,结合姿态变形与运动感知调制。
- 在Occ3D-nuScenes上mIoU提升7.67%,推理速度加快2.84倍。
- 适合自动驾驶长时序场景预测,尤其关注结构几何精度。
4D占用预测建模3D场景随时间演变,对自动驾驶中罕见场景模拟至关重要。现有方法多依赖离散标记化与自回归预测,但在静态结构上易产生几何失真,且长时间预测时时间一致性不足。本文提出几何感知时空建模方法(GAST),基于渐进式显式-隐式生成与双路径时空建模。生成模块通过姿态驱动形变、运动感知特征调制与注意力特征优化,实现每帧高几何保真度与语义合理性。时空模块通过全局上下文聚合增强空间一致性,同时提取时序动态捕捉场景演化。该统一设计支持历史重建与未来预测端到端联合优化。在Occ3D-nuScenes数据集上的实验表明,本方法相比最先进方法,mIoU提升7.67%,IoU提升6.44%,推理速度加快2.84倍,且长期预测性能优异。
原文摘要 · Abstract (English)
4D occupancy forecasting models the spatio-temporal evolution of 3D scenes and is crucial for autonomous driving, especially for corner-case simulation. Existing methods often rely on discrete tokenization followed by autoregressive prediction, yet struggle with geometric distortion in static structures and inconsistent temporal coherence over the forecasting horizon. In this work, we propose a Geometry-Aware Spatio-Temporal context modeling method (GAST) for 4D occupancy forecasting, built upon progressive explicit-implicit generation and dual-path spatio-temporal modeling. Specifically, the generation module produces per-frame occupancy with high geometric fidelity and semantic plausibility through pose-driven warping, motion-aware feature modulation, and attention-based feature refinement. Subsequently, the spatio-temporal module enhances spatial consistency through global context aggregation while capturing scene evolution through temporal dynamics extraction. This unified design enables joint optimization of historical reconstruction and future forecasting in an end-to-end manner. Extensive experiments on Occ3D-nuScenes demonstrate the superiority of our method, outperforming the state-of-the-art by 7.67% in mIoU and 6.44% in IoU with a 2.84x speedup, while maintaining strong performance in long-term forecasting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。