arXiv:2411.10171cs.ROcs.AI2024-11被引 6

用高保真扩散模型提升自动驾驶决策能力,减少试错成本。

Imagine-2-Drive: Leveraging High-Fidelity World Models via Multi-Modal Diffusion Policies

  • 用扩散模型生成多帧未来画面,减少误差累积。
  • 新策略模型让驾驶行为更灵活,成功率提升20%。
  • 适合追求安全高效训练的自动驾驶研究者。

基于世界模型的强化学习(WMRL)通过减少在线交互,实现样本高效的策略学习,尤其适用于成本高且危险的自动驾驶场景。然而,现有世界模型常因预测精度低和逐步误差累积,导致长时程决策性能下降。同时,传统强化学习策略多为确定性或单高斯分布,难以捕捉复杂驾驶场景中的多模态决策特性。为此,我们提出Imagine-2-Drive,一个融合高保真世界模型与多模态扩散策略的新型WMRL框架。其核心包括:DiffDreamer——一种基于扩散的世界模型,可一次性生成多帧未来观测,缓解误差累积;DPA(Diffusion Policy Actor)——一种扩散策略,用于建模多样化、多模态的轨迹分布。在CARLA平台上,通过标准驾驶基准测试验证,该方法相较已有世界模型基线,在路线完成率和成功率上分别提升15%和20%,实现仅需少量在线交互的鲁棒策略学习。

原文摘要 · Abstract (English)

World Model-based Reinforcement Learning (WMRL) enables sample efficient policy learning by reducing the need for online interactions which can potentially be costly and unsafe, especially for autonomous driving. However, existing world models often suffer from low prediction fidelity and compounding one-step errors, leading to policy degradation over long horizons. Additionally, traditional RL policies, often deterministic or single Gaussian-based, fail to capture the multi-modal nature of decision-making in complex driving scenarios. To address these challenges, we propose Imagine-2-Drive, a novel WMRL framework that integrates a high-fidelity world model with a multi-modal diffusion-based policy actor. It consists of two key components: DiffDreamer, a diffusion-based world model that generates future observations simultaneously, mitigating error accumulation, and DPA (Diffusion Policy Actor), a diffusion-based policy that models diverse and multi-modal trajectory distributions. By training DPA within DiffDreamer, our method enables robust policy learning with minimal online interactions. We evaluate our method in CARLA using standard driving benchmarks and demonstrate that it outperforms prior world model baselines, improving Route Completion and Success Rate by 15% and 20% respectively.

自动驾驶扩散模型世界模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。