可控风险的多视角驾驶场景生成,提升自动驾驶测试真实性
Risk-Controllable Multi-View Diffusion for Driving Scenario Generation
- 将风险目标融入物理驱动模型,生成高风险动态轨迹
- 在nuScenes上使3D检测mAP从18.17提升至30.50,FID降至15.70
- 适合自动驾驶系统测试与安全评估,尤其关注罕见高危场景
生成安全关键型驾驶场景对评估和改进自动驾驶系统至关重要,但长尾风险场景在真实数据中罕见,且难以通过人工设计明确。现有生成方法通常将风险作为事后标签,难以保持多视角场景的几何一致性。我们提出RiskMV-DPO,一种物理驱动、风险可控的多视角场景生成系统。通过融合目标风险等级与物理基础风险建模,合成多样且高风险的动态轨迹,作为扩散视频生成器的显式几何锚点。为确保时空一致性和几何保真度,引入几何-外观对齐模块和区域感知直接偏好优化(RA-DPO)策略,结合运动感知掩码聚焦局部动态区域学习。在nuScenes数据集上的实验表明,RiskMV-DPO可自由生成多样化场景,同时保持视觉质量,使3D检测mAP从18.17提升至30.50,FID降低至15.70。本工作将世界模型角色从被动环境预测转变为主动、风险可控的合成,为具身智能开发提供可扩展工具链。
原文摘要 · Abstract (English)
Generating safety-critical driving scenarios is crucial for evaluating and improving autonomous driving systems, but long-tail risky situations are rarely observed in real-world data and difficult to specify through manual scenario design. Existing generative approaches typically treat risk as an after-the-fact label and struggle to maintain geometric consistency in multi-view driving scenes. We present RiskMV-DPO, a general and systematic pipeline for physically-informed, risk-controllable multi-view scenario generation. By integrating target risk levels with physically-grounded risk modeling, we synthesize diverse and high-stakes dynamic trajectories that serve as explicit geometric anchors for a diffusion-based video generator. To ensure spatial-temporal coherence and geometric fidelity, we introduce a geometry-appearance alignment module and a region-aware direct preference optimization (RA-DPO) strategy with motion-aware masking to focus learning on localized dynamic regions. Experiments on the nuScenes dataset show that RiskMV-DPO can freely generate a wide spectrum of diverse scenarios while maintaining visual quality, improving 3D detection mAP from 18.17 to 30.50 and reducing FID to 15.70. Our work shifts the role of world models from passive environment prediction to proactive, risk-controllable synthesis, providing a scalable toolchain for the development of embodied intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。