arXiv:2603.12864cs.CV2026-03被引 3

分离驾驶场景要素,生成更危险的对抗性交通场景

Composing Driving Worlds through Disentangled Control for Adversarial Scenario Generation

  • 分解场景结构、物体身份和驾驶动作,实现独立控制
  • 身份编辑FVD提升17%,动作控制旋转/平移误差降30%和47%
  • 适合自动驾驶系统压力测试,发现规划器严重失效

自动驾驶面临安全关键边缘案例的‘长尾’问题,这些案例常由常见交通元素的异常组合引发。现有可控生成模型存在引导不完整或因素纠缠,难以独立操控场景结构、物体身份与自车动作。我们提出CompoSIA,一种解耦的驾驶视频模拟器,支持对多样化对抗性场景进行细粒度控制。为实现元素身份的可控替换,提出噪声级身份注入机制,仅需一张参考图像即可生成不同姿态下的身份信息。此外,引入分层双分支动作控制机制,提升动作可控性。这种解耦控制可系统性地将安全元素组合成危险配置,是传统纠缠生成器无法实现的。大量对比实验表明,其在身份编辑上的FVD指标优于现有最优方法17%,动作控制的旋转误差减少30%,平移误差减少47%。下游压力测试显示显著规划失败:在各类编辑下,3秒内平均碰撞率上升173%。

原文摘要 · Abstract (English)

A major challenge in autonomous driving is the "long tail" of safety-critical edge cases, which often emerge from unusual combinations of common traffic elements. Synthesizing these scenarios is crucial, yet current controllable generative models provide incomplete or entangled guidance, preventing the independent manipulation of scene structure, object identity, and ego actions. We introduce CompoSIA, a compositional driving video simulator that disentangles these traffic factors, enabling fine-grained control over diverse adversarial driving scenarios. To support controllable identity replacement of scene elements, we propose a noise-level identity injection, allowing pose-agnostic identity generation across diverse element poses, all from a single reference image. Furthermore, a hierarchical dual-branch action control mechanism is introduced to improve action controllability. Such disentangled control enables adversarial scenario synthesis-systematically combining safe elements into dangerous configurations that entangled generators cannot produce. Extensive comparisons demonstrate superior controllable generation quality over state-of-the-art baselines, with a 17% improvement in FVD for identity editing and reductions of 30% and 47% in rotation and translation errors for action control. Furthermore, downstream stress-testing reveals substantial planner failures: across editing modalities, the average collision rate of 3s increases by 173%.

对抗生成自动驾驶场景模拟解耦控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。