无需预设轨迹库,用去相关表示提升自动驾驶轨迹多样性。
TransDiffuser: Diverse Trajectory Generation with Decorrelated Multi-modal Representation for End-to-end Autonomous Driving
- 通过解码器中的多模态表征去相关优化,增强潜在空间多样性。
- 在NAV SIM基准上实现94.85的PDMS,优于现有方法。
- 适合追求高多样性和强泛化能力的端到端自动驾驶研究者。
近年来,扩散模型在视觉生成、语言建模等多个领域展现出巨大潜力。将生成能力迁移至端到端自动驾驶系统也成为重要方向。然而,现有基于扩散的轨迹生成模型常出现模式崩溃问题,不同随机噪声经去噪过程后趋于相似轨迹。为此,当前先进模型依赖预定义轨迹词典或训练集中的场景先验来缓解崩溃并提升多样性,但这些归纳偏置在真实部署中不可用,难以泛化至未见场景。本文提出TransDiffuser,一种编码器-解码器结构的生成式轨迹规划模型,将场景信息与运动状态作为多模态条件输入解码器。不同于以往方法,我们在去噪过程中引入简单有效的多模态表征去相关优化机制,丰富潜在表示空间,更好引导下游生成。无需预定义轨迹锚点或预先计算的场景先验,TransDiffuser在面向闭环规划的基准NAV SIM上取得94.85的PDMS,超越现有最优方法。定性评估表明,该模型生成的轨迹更具多样性且更合理,能探索更大可行驶区域。
原文摘要 · Abstract (English)
In recent years, diffusion models have demonstrated remarkable potential across diverse domains, from vision generation to language modeling. Transferring its generative capabilities to modern end-to-end autonomous driving systems has also emerged as a promising direction. However, existing diffusion-based trajectory generative models often exhibit mode collapse where different random noises converge to similar trajectories after the denoising process.Therefore, state-of-the-art models often rely on anchored trajectories from pre-defined trajectory vocabulary or scene priors in the training set to mitigate collapse and enrich the diversity of generated trajectories, but such inductive bias are not available in real-world deployment, which can be challenged when generalizing to unseen scenarios. In this work, we investigate the possibility of effectively tackling the mode collapse challenge without the assumption of pre-defined trajectory vocabulary or pre-computed scene priors. Specifically, we propose TransDiffuser, an encoder-decoder based generative trajectory planning model, where the encoded scene information and motion states serve as the multi-modal conditional input of the denoising decoder. Different from existing approaches, we exploit a simple yet effective multi-modal representation decorrelation optimization mechanism during the denoising process to enrich the latent representation space which better guides the downstream generation. Without any predefined trajectory anchors or pre-computed scene priors, TransDiffuser achieves the PDMS of 94.85 on the closed-loop planning-oriented benchmark NAVSIM, surpassing previous state-of-the-art methods. Qualitative evaluation further showcases TransDiffuser generates more diverse and plausible trajectories which explore more drivable area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。