探索视频生成与世界模型如何协同提升自动驾驶安全性
Exploring the Interplay Between Video Generation and World Models in Autonomous Driving: A Survey
- 基于扩散模型的结构相似性,融合视频生成与世界模型
- 对比JEPA、Genie、Sora等模型,揭示设计多样性
- 适合关注自动驾驶仿真与决策优化的研究者
世界模型与视频生成是自动驾驶领域的关键技术,分别用于模拟真实环境动态和生成逼真视频序列。二者正逐步融合,以增强车辆的情境感知与决策能力。本文聚焦于两者在扩散模型结构上的共性,探讨其如何提升驾驶场景模拟的准确性与连贯性。研究分析了JEPA、Genie、Sora等代表性工作,展现世界模型设计的多元路径,凸显该领域尚无统一定义的现状。同时,论文讨论了关键评估指标,如3D场景重建的Chamfer距离与视频质量评估的Fréchet Inception Distance(FID)。通过剖析两者的交互关系,本文识别出当前核心挑战与未来方向,强调二者协同对推动自动驾驶系统更安全、可靠发展的潜力。
原文摘要 · Abstract (English)
World models and video generation are pivotal technologies in the domain of autonomous driving, each playing a critical role in enhancing the robustness and reliability of autonomous systems. World models, which simulate the dynamics of real-world environments, and video generation models, which produce realistic video sequences, are increasingly being integrated to improve situational awareness and decision-making capabilities in autonomous vehicles. This paper investigates the relationship between these two technologies, focusing on how their structural parallels, particularly in diffusion-based models, contribute to more accurate and coherent simulations of driving scenarios. We examine leading works such as JEPA, Genie, and Sora, which exemplify different approaches to world model design, thereby highlighting the lack of a universally accepted definition of world models. These diverse interpretations underscore the field's evolving understanding of how world models can be optimized for various autonomous driving tasks. Furthermore, this paper discusses the key evaluation metrics employed in this domain, such as Chamfer distance for 3D scene reconstruction and Fréchet Inception Distance (FID) for assessing the quality of generated video content. By analyzing the interplay between video generation and world models, this survey identifies critical challenges and future research directions, emphasizing the potential of these technologies to jointly advance the performance of autonomous driving systems. The findings presented in this paper aim to provide a comprehensive understanding of how the integration of video generation and world models can drive innovation in the development of safer and more reliable autonomous vehicles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。