arXiv:2412.13772cs.CV2024-12被引 20

用解耦动态流与图像辅助训练,提升3D占用率世界模型效率

An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training

  • 将占据预测拆解为动态体素变形与静态体素变换,简化训练流程
  • 在nuScenes和OpenScene上达到顶尖4D场景预测性能,计算成本更低
  • 结合可微体渲染生成深度图,通过光照一致性增强预测可靠性

自动驾驶领域对世界模型的兴趣日益增长,旨在基于历史观测预测未来潜在场景。本文提出DFIT-OccWorld,一种高效的3D占据世界模型,采用解耦动态流与图像辅助训练策略,显著提升4D场景预测性能。为简化训练过程,摒弃以往两阶段训练方式,创新性地将占据预测重构为解耦体素变形过程:通过体素流对现有观测进行变形以预测未来动态体素,而静态体素则通过位姿变换获得。此外,引入图像辅助训练范式以增强预测可靠性,具体采用可微体渲染生成预测未来体积对应的渲染深度图,并用于基于渲染的光度一致性约束。实验表明,该方法在nuScenes和OpenScene基准上均实现顶尖的4D占据预测、端到端运动规划及点云预测性能,相比现有3D世界模型表现更优,且计算开销显著降低。

原文摘要 · Abstract (English)

The field of autonomous driving is experiencing a surge of interest in world models, which aim to predict potential future scenarios based on historical observations. In this paper, we introduce DFIT-OccWorld, an efficient 3D occupancy world model that leverages decoupled dynamic flow and image-assisted training strategy, substantially improving 4D scene forecasting performance. To simplify the training process, we discard the previous two-stage training strategy and innovatively reformulate the occupancy forecasting problem as a decoupled voxels warping process. Our model forecasts future dynamic voxels by warping existing observations using voxel flow, whereas static voxels are easily obtained through pose transformation. Moreover, our method incorporates an image-assisted training paradigm to enhance prediction reliability. Specifically, differentiable volume rendering is adopted to generate rendered depth maps through predicted future volumes, which are adopted in render-based photometric consistency. Experiments demonstrate the effectiveness of our approach, showcasing its state-of-the-art performance on the nuScenes and OpenScene benchmarks for 4D occupancy forecasting, end-to-end motion planning and point cloud forecasting. Concretely, it achieves state-of-the-art performances compared to existing 3D world models while incurring substantially lower computational costs.

3D占据世界模型自动驾驶扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。