用时空记忆提升自动驾驶3D占位预测的时序一致性
Occupancy Learning with Spatiotemporal Memory
- 引入场景级时空记忆,高效存储历史信息
- 提升3D占位预测性能,mIoU提升3,时序不一致降低29%
- 适合需要高精度动态环境建模的自动驾驶研究
3D占位成为自动驾驶精细环境感知的有前景表示方法。然而,由于处理成本高以及体素的不确定性和动态性,跨多帧高效聚合3D占位仍具挑战。为此,我们提出ST-Occ,一种场景级占位表示学习框架,能有效学习具有时间一致性的时空特征。ST-Occ包含两项核心设计:一是时空记忆,通过场景级表示高效捕获并存储全面的历史信息;二是记忆注意力,基于时空记忆对当前占位表示进行条件建模,具备不确定性与动态感知能力。该方法通过利用多帧输入间的时序依赖,显著增强3D占位预测任务的时空表征。实验表明,本方法在性能上超越当前最优方法3 mIoU,同时将时序不一致性降低29%。
原文摘要 · Abstract (English)
3D occupancy becomes a promising perception representation for autonomous driving to model the surrounding environment at a fine-grained scale. However, it remains challenging to efficiently aggregate 3D occupancy over time across multiple input frames due to the high processing cost and the uncertainty and dynamics of voxels. To address this issue, we propose ST-Occ, a scene-level occupancy representation learning framework that effectively learns the spatiotemporal feature with temporal consistency. ST-Occ consists of two core designs: a spatiotemporal memory that captures comprehensive historical information and stores it efficiently through a scene-level representation and a memory attention that conditions the current occupancy representation on the spatiotemporal memory with a model of uncertainty and dynamic awareness. Our method significantly enhances the spatiotemporal representation learned for 3D occupancy prediction tasks by exploiting the temporal dependency between multi-frame inputs. Experiments show that our approach outperforms the state-of-the-art methods by a margin of 3 mIoU and reduces the temporal inconsistency by 29%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。