用任意轨迹和车辆编辑逼真驾驶场景,一次生成多样变化。
HorizonForge: Driving Scene Editing with Any Trajectories and Any Vehicles
- 将场景重构为可编辑的高斯点云与网格,支持精细3D操作
- 通过噪声感知视频扩散生成一致时空画面,单次前向传播完成
- 适合自动驾驶仿真、影视场景生成等需要可控逼真场景的领域
可控驾驶场景生成对实现真实且可扩展的自动驾驶仿真至关重要,但现有方法难以同时兼顾视觉逼真度与精确控制。我们提出 HorizonForge,一个统一框架,将场景重建为可编辑的高斯点云与网格,支持细粒度3D操作及语言驱动的车辆插入。通过噪声感知视频扩散过程渲染修改内容,保证时空一致性,可在单次前向传播中生成多样化场景变化,无需逐轨迹优化。为标准化评估,我们进一步构建 HorizonSuite,一个涵盖自车与智能体级编辑任务(如轨迹调整、物体操作)的综合性基准。大量实验表明,高斯-网格表示显著优于其他3D表示,视频扩散中的时序先验对合成连贯性至关重要。结合这些发现,HorizonForge 建立了一种简单而强大的逼真可控驾驶仿真范式,在用户偏好上比次优方法提升83.4%,FID指标改善25.19%。
原文摘要 · Abstract (English)
Controllable driving scene generation is critical for realistic and scalable autonomous driving simulation, yet existing approaches struggle to jointly achieve photorealism and precise control. We introduce HorizonForge, a unified framework that reconstructs scenes as editable Gaussian Splats and Meshes, enabling fine-grained 3D manipulation and language-driven vehicle insertion. Edits are rendered through a noise-aware video diffusion process that enforces spatial and temporal consistency, producing diverse scene variations in a single feed-forward pass without per-trajectory optimization. To standardize evaluation, we further propose HorizonSuite, a comprehensive benchmark spanning ego- and agent-level editing tasks such as trajectory modifications and object manipulation. Extensive experiments show that Gaussian-Mesh representation delivers substantially higher fidelity than alternative 3D representations, and that temporal priors from video diffusion are essential for coherent synthesis. Combining these findings, HorizonForge establishes a simple yet powerful paradigm for photorealistic, controllable driving simulation, achieving an 83.4% user-preference gain and a 25.19% FID improvement over the second best state-of-the-art method. Project page: https://horizonforge.github.io/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。