arXiv:2502.07309cs.CV2025-02ICLR被引 33

用2D标签训练3D占位模型,降低自动驾驶标注成本

Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving

  • 两阶段训练:先自监督预训练,再微调
  • 仅需2D标签即可实现4D场景预测
  • 适合数据标注受限的自动驾驶研究

理解世界动态对自动驾驶规划至关重要。现有方法通过学习3D占位世界模型来预测未来环境,但依赖昂贵的3D占位标注。针对室外3D标注成本高的问题,本文提出半监督视觉中心的3D占位世界模型PreWorld,通过创新的两阶段训练范式,利用2D标签挖掘潜力:预训练阶段采用属性投影头生成场景的RGB、密度、语义等属性场,借助体渲染技术从2D标签获得时序监督;进一步引入状态条件预测模块,直接递归预测未来占位与自车轨迹。在nuScenes数据集上的大量实验验证了方法的有效性与可扩展性,PreWorld在3D占位预测、4D占位预测和运动规划任务中均达到有竞争力的表现。

原文摘要 · Abstract (English)

Understanding world dynamics is crucial for planning in autonomous driving. Recent methods attempt to achieve this by learning a 3D occupancy world model that forecasts future surrounding scenes based on current observation. However, 3D occupancy labels are still required to produce promising results. Considering the high annotation cost for 3D outdoor scenes, we propose a semi-supervised vision-centric 3D occupancy world model, PreWorld, to leverage the potential of 2D labels through a novel two-stage training paradigm: the self-supervised pre-training stage and the fully-supervised fine-tuning stage. Specifically, during the pre-training stage, we utilize an attribute projection head to generate different attribute fields of a scene (e.g., RGB, density, semantic), thus enabling temporal supervision from 2D labels via volume rendering techniques. Furthermore, we introduce a simple yet effective state-conditioned forecasting module to recursively forecast future occupancy and ego trajectory in a direct manner. Extensive experiments on the nuScenes dataset validate the effectiveness and scalability of our method, and demonstrate that PreWorld achieves competitive performance across 3D occupancy prediction, 4D occupancy forecasting and motion planning tasks.

3D占位半监督自动驾驶视觉中心

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。