arXiv:2602.23894cs.CV2026-02中稿 · version

无需标注数据,端到端自监督预测动态场景的3D占据与运动流。

SelfOccFlow: Towards end-to-end self-supervised 3D Occupancy Flow prediction

  • 将场景分解为静态与动态的符号距离场,通过时间聚合隐式学习运动。
  • 在SemanticKITTI等数据集上达到优于现有方法的占据与运动估计精度。
  • 适合自动驾驶中动态环境感知,尤其适用于无标注数据场景。

估计车辆周围环境的3D占据与运动对于自动驾驶至关重要,有助于在动态环境中实现态势感知。现有方法虽联合学习几何与运动,但依赖昂贵的3D占据和运动标注、边界框中的速度标签或预训练光流模型。我们提出一种自监督的3D占据流估计方法,无需人工标注或外部光流监督。该方法将场景解耦为独立的静态与动态符号距离场,并通过时间聚合隐式学习运动。此外,我们引入一种基于特征余弦相似性的强自监督流线索。我们在SemanticKITTI、KITTI-MOT和nuScenes数据集上验证了所提方法的有效性。

原文摘要 · Abstract (English)

Estimating 3D occupancy and motion at the vehicle's surroundings is essential for autonomous driving, enabling situational awareness in dynamic environments. Existing approaches jointly learn geometry and motion but rely on expensive 3D occupancy and flow annotations, velocity labels from bounding boxes, or pretrained optical flow models. We propose a self-supervised method for 3D occupancy flow estimation that eliminates the need for human-produced annotations or external flow supervision. Our method disentangles the scene into separate static and dynamic signed distance fields and learns motion implicitly through temporal aggregation. Additionally, we introduce a strong self-supervised flow cue derived from features' cosine similarities. We demonstrate the efficacy of our 3D occupancy flow method on SemanticKITTI, KITTI-MOT, and nuScenes.

3D占据自监督自动驾驶运动估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。