用偏好优化提升扩散模型生成更真实可控的交通场景
Direct Preference Optimization-Enhanced Multi-Guided Diffusion Model for Traffic Scenario Generation
- 多任务框架支持多种引导输入,保持真实交通先验
- 采用DPO算法微调,实现高精度引导采样
- 在nuScenes数据集上兼顾真实、多样与可控性
基于扩散的模型能有效利用真实驾驶数据生成逼真且多样的交通场景。这些模型通过引导采样融入特定交通偏好以提升场景真实性。然而,引导采样过程可能导致偏离真实交通先验,引发不合理的交通行为。为此,我们提出一种多引导扩散模型,采用新型训练策略,在使用多种引导组合时仍能紧密遵循交通先验。该模型采用多任务学习框架,使单一扩散模型可处理多种引导输入。为提升引导采样精度,模型通过直接偏好优化(DPO)算法进行微调,基于引导评分优化偏好,有效应对引导采样微调过程中昂贵且常不可导的梯度计算难题。在nuScenes数据集上的评估表明,该模型在真实度、多样性和可控性之间提供了强有力的基准。
原文摘要 · Abstract (English)
Diffusion-based models are recognized for their effectiveness in using real-world driving data to generate realistic and diverse traffic scenarios. These models employ guided sampling to incorporate specific traffic preferences and enhance scenario realism. However, guiding the sampling process to conform to traffic rules and preferences can result in deviations from real-world traffic priors and potentially leading to unrealistic behaviors. To address this challenge, we introduce a multi-guided diffusion model that utilizes a novel training strategy to closely adhere to traffic priors, even when employing various combinations of guides. This model adopts a multi-task learning framework, enabling a single diffusion model to process various guide inputs. For increased guided sampling precision, our model is fine-tuned using the Direct Preference Optimization (DPO) algorithm. This algorithm optimizes preferences based on guide scores, effectively navigating the complexities and challenges associated with the expensive and often non-differentiable gradient calculations during the guided sampling fine-tuning process. Evaluated using the nuScenes dataset our model provides a strong baseline for balancing realism, diversity and controllability in the traffic scenario generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。