提出帧级自回归视频生成框架FAR-Drive,实现自动驾驶闭环仿真。
FAR-Drive: Frame-AutoRegressive Video Generation in Closed-Loop Autonomous Driving
- 采用多视角扩散变换器与细粒度结构控制,保证多摄像头几何一致性。
- 两阶段训练策略提升长时序一致性和自条件迭代下的稳定性。
- 单卡推理延迟低于1秒,适合实时交互式自动驾驶仿真。
尽管自动驾驶技术进展迅速,但可靠训练与评估仍受限于缺乏可扩展且交互式的仿真环境。现有生成视频模型虽视觉保真度高,但多为开环运行,无法支持智能体动作与环境演化的帧级精细交互。构建基于学习的自动驾驶闭环仿真面临三大挑战:长时序跨视角一致性维持、迭代自条件下的自回归退化抑制、低延迟推理约束满足。本文提出FAR-Drive,一种面向自动驾驶的帧级自回归视频生成框架。引入多视角扩散变换器与细粒度结构控制,实现多相机几何一致性生成。针对长时序一致性与迭代退化问题,设计包含自适应参考时域条件与混合强迫自回归训练的两阶段训练策略,逐步提升自条件下的一致性与鲁棒性。为满足低延迟交互需求,进一步集成系统级效率优化以加速推理。在nuScenes数据集上的实验表明,该方法在现有闭环自动驾驶仿真中达到领先性能,同时在单张GPU上保持亚秒级延迟。
原文摘要 · Abstract (English)
Despite rapid progress in autonomous driving, reliable training and evaluation of driving systems remain fundamentally constrained by the lack of scalable and interactive simulation environments. Recent generative video models achieve remarkable visual fidelity, yet most operate in open-loop settings and fail to support fine-grained frame-level interaction between agent actions and environment evolution. Building a learning-based closed-loop simulator for autonomous driving poses three major challenges: maintaining long-horizon temporal and cross-view consistency, mitigating autoregressive degradation under iterative self-conditioning, and satisfying low-latency inference constraints. In this work, we propose FAR-Drive, a frame-level autoregressive video generation framework for autonomous driving. We introduce a multi-view diffusion transformer with fine-grained structured control, enabling geometrically consistent multi-camera generation. To address long-horizon consistency and iterative degradation, we design a two-stage training strategy consisting of adaptive reference horizon conditioning and blend-forcing autoregressive training, which progressively improves consistency and robustness under self-conditioning. To meet low-latency interaction requirements, we further integrate system-level efficiency optimizations for inference acceleration. Experiments on the nuScenes dataset demonstrate that our method achieves state-of-the-art performance among existing closed-loop autonomous driving simulation approaches, while maintaining sub-second latency on a single GPU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。