用视觉输入生成可控又逼真的交通场景,提升自动驾驶测试效果。
End-to-end Conditional Diffusion for Realistic and Controllable Visual Traffic Scenario Generation

- 端到端扩散模型联合生成车辆状态与控制指令
- 在Bench2Drive上实现可控性与真实性的良好平衡
- 适合自动驾驶系统测试,尤其安全高风险场景
生成既真实又可控的闭环交通场景对评估自动驾驶系统至关重要,尤其在罕见危险交互场景下。现有基于学习的方法难以兼顾可控性与真实性,或控制粒度不足,或牺牲行为合理性。本文提出E2E-CDiff,一种基于前视视觉观测的端到端条件扩散框架,联合去噪未来运动状态与可执行的低层控制指令,用于路线交互背景车辆。该统一的状态-动作生成机制缓解了传统两阶段轨迹-控制器流程中的规划-控制不匹配问题。可微分引导机制进一步调节速度、强制符合可行驶区域,并支持避撞或主动碰撞行为,实现自然与安全关键场景生成。在Bench2Drive上的实验表明,相比代表性强化与模仿学习基线,E2E-CDiff在可控性与真实性间取得更优权衡;其碰撞引导变体可在多个自动驾驶系统中诱发复杂交互。此外,E2E-CDiff作为学习型主车规划器也表现优异,证明了端到端状态-动作扩散的通用性。
原文摘要 · Abstract (English)
Generating closed-loop traffic scenarios that are both realistic and controllable is crucial for evaluating autonomous driving systems, especially under rare safety-critical interactions. Existing learning-based methods often struggle to balance controllability and realism, offering either limited fine-grained control over traffic behavior or controllable scenarios at the expense of behavioral plausibility. This paper presents E2E-CDiff, an end-to-end conditional diffusion framework for controllable and realistic scenario generation. Conditioned on front-view visual observations, E2E-CDiff jointly denoises future motion states and executable low-level controls for route-interacting background vehicles. This unified state-action generation mitigates the planning-control mismatch in conventional two-stage trajectory-then-controller pipelines. Differentiable guidance further regulates speed, enforces drivable-area compliance, and supports collision-avoidance or collision-seeking behaviors, enabling both naturalistic and safety-critical scenario generation. Experiments on Bench2Drive show that E2E-CDiff achieves a favorable controllability-realism trade-off compared with representative reinforcement- and imitation-learning baselines, while its collision-guided variant induces challenging interactions across multiple autonomous driving systems. E2E-CDiff also performs competitively as a learning-based ego planner, demonstrating the generality of end-to-end state-action diffusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。