arXiv:2410.15461cs.CVcs.MM2024-10被引 25

提出EVA世界模型,实现更精准的多步视频预测。

EVA: An Embodied World Model for Future Video Anticipation

  • 结合视觉语言与视频生成模型,设计中间推理策略提升预测能力。
  • 在多任务基准上表现优异,支持长序列自回归生成。
  • 适合机器人、自动驾驶等需要未来视频预判的场景。

视频生成模型在模拟未来状态方面取得显著进展,展现出作为具身场景中世界模型的潜力。然而,现有模型普遍缺乏稳健理解,难以进行多步预测或处理分布外(OOD)场景。为此,我们提出反射生成(RoG)策略,利用预训练视觉语言模型与视频生成模型的互补优势,使它们能在具身场景中充当世界模型。为支持RoG,我们构建了具身视频预判基准(EVA-Bench),涵盖多种任务和场景,使用域内与域外数据集进行评估。在此基础上,我们设计了具身视频预判器(EVA),采用多阶段训练范式生成高保真视频帧,并通过自回归策略实现对更长视频序列的自适应泛化。大量实验表明,EVA在视频生成与机器人等下游任务中表现出色,为大规模预训练模型在真实世界视频预测中的应用铺平道路。视频演示见:https://sites.google.com/view/icml-eva

原文摘要 · Abstract (English)

Video generation models have made significant progress in simulating future states, showcasing their potential as world simulators in embodied scenarios. However, existing models often lack robust understanding, limiting their ability to perform multi-step predictions or handle Out-of-Distribution (OOD) scenarios. To address this challenge, we propose the Reflection of Generation (RoG), a set of intermediate reasoning strategies designed to enhance video prediction. It leverages the complementary strengths of pre-trained vision-language and video generation models, enabling them to function as a world model in embodied scenarios. To support RoG, we introduce Embodied Video Anticipation Benchmark(EVA-Bench), a comprehensive benchmark that evaluates embodied world models across diverse tasks and scenarios, utilizing both in-domain and OOD datasets. Building on this foundation, we devise a world model, Embodied Video Anticipator (EVA), that follows a multistage training paradigm to generate high-fidelity video frames and apply an autoregressive strategy to enable adaptive generalization for longer video sequences. Extensive experiments demonstrate the efficacy of EVA in various downstream tasks like video generation and robotics, thereby paving the way for large-scale pre-trained models in real-world video prediction applications. The video demos are available at \hyperlink{https://sites.google.com/view/icml-eva}{https://sites.google.com/view/icml-eva}.

视频预测世界模型具身智能自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。