arXiv:2603.21104cs.ROcs.CV2026-03被引 2

通过因果推理生成更真实的高危驾驶场景,提升评估安全性。

CounterScene: Counterfactual Causal Reasoning in Generative World Models for Safety-Critical Closed-Loop Evaluation

  • 基于因果关系识别关键冲突主体并建模动态交互依赖。
  • 长时序下碰撞率提升至22.7%,轨迹真实度优于基线(ADE 1.88 vs. 2.09)。
  • 适用于自动驾驶安全评估,尤其适合需要高真实感的闭环测试。

生成高危驾驶场景需理解危险交互的成因,而非简单制造碰撞。现有方法依赖启发式对抗代理选择和无结构扰动,缺乏对交互依赖的显式建模,导致真实度与对抗性之间存在权衡。我们提出CounterScene,为闭环生成式鸟瞰图世界模型赋予结构化反事实推理能力。给定安全场景,其核心问题为:若关键因果代理行为不同会怎样?为此,我们引入因果对抗代理识别以定位关键代理并分类冲突类型,并构建冲突感知的交互式世界模型,利用因果交互图显式建模动态多智能体依赖关系。在此基础上,阶段自适应反事实引导对识别出的代理进行最小干预,移除其时空安全裕度,使风险通过自然交互传播而显现。在nuScenes上的大量实验表明,CounterScene在保持最优轨迹真实度的同时实现了最强对抗效果,长时序碰撞率从12.3%提升至22.7%,相比最强基线真实度更高(ADE 1.88 vs. 2.09)。该优势在更长滚动窗口下进一步扩大,且零样本泛化至nuPlan,达到当前最佳真实度表现。

原文摘要 · Abstract (English)

Generating safety-critical driving scenarios requires understanding why dangerous interactions arise, rather than merely forcing collisions. However, existing methods rely on heuristic adversarial agent selection and unstructured perturbations, lacking explicit modeling of interaction dependencies and thus exhibiting a realism--adversarial trade-off. We present CounterScene, a framework that endows closed-loop generative BEV world models with structured counterfactual reasoning for safety-critical scenario generation. Given a safe scene, CounterScene asks: what if the causally critical agent had behaved differently? To answer this, we introduce causal adversarial agent identification to identify the critical agent and classify conflict types, and develop a conflict-aware interactive world model in which a causal interaction graph is used to explicitly model dynamic inter-agent dependencies. Building on this structure, stage-adaptive counterfactual guidance performs minimal interventions on the identified agent, removing its spatial and temporal safety margins while allowing risk to emerge through natural interaction propagation. Extensive experiments on nuScenes demonstrate that CounterScene achieves the strongest adversarial effectiveness while maintaining superior trajectory realism across all horizons, improving long-horizon collision rate from 12.3% to 22.7% over the strongest baseline with better realism (ADE 1.88 vs.2.09). Notably, this advantage further widens over longer rollouts, and CounterScene generalizes zero-shot to nuPlan with state-of-the-art realism.

自动驾驶反事实推理生成模型安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。