arXiv:2412.03758cs.CV2024-12被引 1

用语义与图像交替生成,让大模型更连贯地续写自动驾驶视频。

ARCON: Advancing Auto-Regressive Continuation for Driving Videos

  • 交替生成语义图和RGB像素,显式学习视频结构。
  • 无需特殊设计即可保持图像与语义一致性。
  • 结合光流纹理拼接,提升生成视频质量,适合长序列生成。

近期自回归大语言模型的发展推动了其在视频生成中的应用。本文探索利用大视觉模型(LVMs)进行视频续写,该任务对构建世界模型和预测未来帧至关重要。我们提出ARCON,一种交替生成语义与RGB令牌的方案,使LVM能够显式学习高层级视频结构信息。实验发现,无需特殊设计即可实现生成图像与语义图的高度一致。此外,采用基于光流的纹理拼接方法进一步提升视觉质量。在自动驾驶场景下的实验表明,该模型能持续生成长时序视频。

原文摘要 · Abstract (English)

Recent advancements in auto-regressive large language models (LLMs) have led to their application in video generation. This paper explores the use of Large Vision Models (LVMs) for video continuation, a task essential for building world models and predicting future frames. We introduce ARCON, a scheme that alternates between generating semantic and RGB tokens, allowing the LVM to explicitly learn high-level structural video information. We find high consistency in the RGB images and semantic maps generated without special design. Moreover, we employ an optical flow-based texture stitching method to enhance visual quality. Experiments in autonomous driving scenarios show that our model can consistently generate long videos.

视频生成自回归大视觉模型自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。