用视频生成统一建模未来场景,端到端驾驶更高效
SUV: Future Scene Understanding as Video Generation for End-to-End Driving

- 将未来场景理解转为视频生成任务,共享视觉专家
- 仅用单前视摄像头,实现91.0 EPDMS的轨迹预测性能
- 无需候选轨迹选择,适合真实自动驾驶部署
端到端驾驶需要对未来的场景有连贯理解,但现有方法依赖特定任务头和输出格式,可扩展性差。能否用视频生成作为统一预测器?我们提出SUV,一种统一的端到端驾驶框架,将未来场景理解建模为视频生成任务,采用预训练视频基础模型。SUV将未来外观、语义、相对深度及实例级动态建模为视频流,由共享视频专家处理,无需针对不同流设计独立视觉预测头。通过联合视频-动作注意力机制,动作专家关注所有未来流的潜在表示并生成自车轨迹。实验表明,SUV可直接预测全部四类未来流,受控消融实验显示结构化未来监督与直接未来流访问能显著提升轨迹规划得分。仅使用单个前向摄像头且无候选轨迹选择,SUV在NAVSIM-v2两个测试集上分别取得91.0 EPDMS(navtest)和36.9(navhard)的性能,在长尾WOD-E2E基准上达到7.94的竞争力RFS分数。
原文摘要 · Abstract (English)
End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output formats, with limited scalability. Can video generation instead provide a shared predictor? We introduce SUV, a unified end-to-end driving framework that casts future Scene Understanding as Video generation using a pretrained video foundation model. SUV models future appearance, semantics, relative depth, and instance-level dynamics as video streams with a shared video expert, without stream-specific visual prediction heads. Through joint video-action attention, the action expert attends to the latent representations of all future streams and generates the ego trajectory. Experiments show that SUV directly predicts all four future streams, while controlled ablations show that structured future supervision and direct future-stream access yield higher trajectory planning scores. With only a single front camera and no candidate-trajectory selection, SUV outperforms a broad set of recent state-of-the-art methods on both NAVSIM-v2 splits, achieving 91.0 EPDMS on navtest and 36.9 on navhard. On the long-tail WOD-E2E benchmark, SUV achieves a competitive RFS of 7.94.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。