arXiv:2601.01528cs.CVcs.AI2026-01被引 16

首个面向自动驾驶生成式世界模型的综合评测基准,解决真实场景模拟难题。

DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving

  • 构建多源融合数据集,覆盖复杂天气、时段与地理环境
  • 提出视觉真实度、轨迹合理性、时序一致性等四项新指标
  • 揭示通用模型与专用模型在画质与物理合理性间的权衡

视频生成模型作为世界模型的一种,已成为AI前沿热点,有望让智能体通过建模复杂场景的时间演化来预知未来。在自动驾驶领域,这催生了驾驶世界模型:能够预测自身及其它交通参与者未来状态的生成式模拟器,支持可扩展仿真、安全测试边缘案例和生成丰富合成数据。然而,尽管研究活跃,该领域仍缺乏严谨的评估基准。现有评价方法存在局限:通用视频指标忽略关键安全因素;轨迹合理性极少量化;时序与多主体一致性被忽视;对自车条件控制能力未予考量。此外,现有数据集无法覆盖实际部署所需多样性。为此,我们提出DrivingGen,首个针对生成式驾驶世界模型的综合性评测基准。DrivingGen结合来自驾驶数据集与互联网级视频源的多样化评估数据集,涵盖不同天气、时段、地理区域与复杂操作;并配套一套新指标,联合评估视觉真实度、轨迹合理性、时序连贯性与可控性。对14个前沿模型的基准测试揭示显著权衡:通用模型视觉效果更优但违背物理规律,驾驶专用模型运动更真实但画质较差。DrivingGen提供统一评估框架,推动可信赖、可控且可部署的驾驶世界模型发展,支撑大规模仿真、规划与数据驱动决策。

原文摘要 · Abstract (English)

Video generation models, as one form of world models, have emerged as one of the most exciting frontiers in AI, promising agents the ability to imagine the future by modeling the temporal evolution of complex scenes. In autonomous driving, this vision gives rise to driving world models: generative simulators that imagine ego and agent futures, enabling scalable simulation, safe testing of corner cases, and rich synthetic data generation. Yet, despite fast-growing research activity, the field lacks a rigorous benchmark to measure progress and guide priorities. Existing evaluations remain limited: generic video metrics overlook safety-critical imaging factors; trajectory plausibility is rarely quantified; temporal and agent-level consistency is neglected; and controllability with respect to ego conditioning is ignored. Moreover, current datasets fail to cover the diversity of conditions required for real-world deployment. To address these gaps, we present DrivingGen, the first comprehensive benchmark for generative driving world models. DrivingGen combines a diverse evaluation dataset curated from both driving datasets and internet-scale video sources, spanning varied weather, time of day, geographic regions, and complex maneuvers, with a suite of new metrics that jointly assess visual realism, trajectory plausibility, temporal coherence, and controllability. Benchmarking 14 state-of-the-art models reveals clear trade-offs: general models look better but break physics, while driving-specific ones capture motion realistically but lag in visual quality. DrivingGen offers a unified evaluation framework to foster reliable, controllable, and deployable driving world models, enabling scalable simulation, planning, and data-driven decision-making.

自动驾驶生成模型仿真评测世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。