用模块化设计生成逼真自动驾驶极端场景,兼顾语义合理与物理真实。
CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis

- 分拆视觉、推理与物理控制三模块,协同生成极端场景。
- 在Waymo数据集上生成符合语义意图且物理可行的逼真视频。
- 适合自动驾驶安全评估与仿真系统研发人员使用。
自动驾驶安全性评估依赖于罕见但高危的交互场景,亟需能主动合成具有逼真视觉观测的极端场景的仿真器。极端场景生成涉及视觉表征、场景推理与车辆轨迹生成控制等多个层面。现有基于先验或模型的方法通常孤立处理场景或轨迹,而扩散方法虽尝试端到端生成,仍难以保证时空一致性与物理真实性。为此,我们提出CARLA-GS,一个将视觉表征、语义推理与物理执行解耦但紧密耦合的模块化极端场景合成框架。从真实驾驶数据出发,重建带几何一致性约束的可编辑高斯场景;多智能体大模型进行场景级推理,识别危险交互并生成意图级航点轨迹;底层运动控制交由CARLA中的PID控制器确保运动学与动力学可行性;最终将模拟车辆状态重投影至高斯场景实现本车视角渲染。该设计实现了高层语义推理、低层物理可执行运动与逼真极端场景生成的统一。在Waymo Open Dataset上的实验表明,该框架能可控生成逼真、时空一致、语义对齐且物理可行的视频。
原文摘要 · Abstract (English)
Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately synthesize corner cases with photorealistic observations. Corner-case generation is inherently a multi-source problem spanning visual representation, scene reasoning, and vehicle trajectory generation and control. Prior knowledge- and model-based approaches typically focus on scene or trajectory components in isolation, while diffusion-based methods attempt end-to-end generation but still struggle to ensure spatiotemporal consistency and physical realism. To unify these aspects within a single framework, we propose CARLA-GS, a modular corner-case synthesis pipeline that decouples visual representation, semantic reasoning, and physics-based execution while maintaining tight cross-module coupling. Starting from real driving data, we reconstruct an editable gaussian scene with additional geometry-consistent constraints. A multi-agent LLM then performs scene-level reasoning to identify risky interactions and generate intent-level waypoint trajectories, while the low-level motion control is delegated to CARLA, where a PID controller ensures kinematic and dynamic feasibility. The simulated vehicle states are finally re-projected into the gaussian scene for ego-centric rendering. This design enables high-level semantic reasoning, low-level physically executable motion, and photorealistic corner-case generation within a unified pipeline. Experiments on the Waymo Open Dataset show, both quantitatively and qualitatively, that our framework enables controllable corner-case generation and produces photorealistic, spatiotemporally consistent videos aligned with semantic intent and physically feasible motion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。