arXiv:2505.20827cs.CV2025-05被引 2

用帧级字幕指导长视频生成,解决多场景故事的误差累积问题。

Frame-Level Captions for Long Video Generation with Complex Multi Scenes

  • 通过帧级标注与注意力机制,实现文本与视频精准对齐。
  • 在VBench 2.0复杂剧情与场景测试中表现优于基线模型。
  • 适合需要多事件、多场景长视频生成的研究者使用。

生成能展现复杂叙事(如剧本改编电影)的长视频具有巨大潜力,远超短片段。然而,当前基于扩散模型的自回归方法常因逐步生成导致严重误差积累(漂移)。此外,多数现有方法聚焦单一连续场景,难以应对多事件、多变化的故事。本文提出新方法:首先构建帧级标注数据集,为复杂多场景长视频提供精细文本引导;结合帧级注意力机制,使每个时间窗口内的帧可接受独立文本提示。训练采用扩散强制(Diffusion Forcing),提升模型对时间的灵活性。在基于WanX2.1-T2V-1.3B模型的VBench 2.0基准测试(“Complex Plots”与“Complex Landscapes”)中,本方法显著提升复杂场景下的指令遵循能力,并生成高质量长视频。相关数据标注方法与训练模型将公开共享。

原文摘要 · Abstract (English)

Generating long videos that can show complex stories, like movie scenes from scripts, has great promise and offers much more than short clips. However, current methods that use autoregression with diffusion models often struggle because their step-by-step process naturally leads to a serious error accumulation (drift). Also, many existing ways to make long videos focus on single, continuous scenes, making them less useful for stories with many events and changes. This paper introduces a new approach to solve these problems. First, we propose a novel way to annotate datasets at the frame-level, providing detailed text guidance needed for making complex, multi-scene long videos. This detailed guidance works with a Frame-Level Attention Mechanism to make sure text and video match precisely. A key feature is that each part (frame) within these windows can be guided by its own distinct text prompt. Our training uses Diffusion Forcing to provide the model with the ability to handle time flexibly. We tested our approach on difficult VBench 2.0 benchmarks ("Complex Plots" and "Complex Landscapes") based on the WanX2.1-T2V-1.3B model. The results show our method is better at following instructions in complex, changing scenes and creates high-quality long videos. We plan to share our dataset annotation methods and trained models with the research community. Project page: https://zgctroy.github.io/frame-level-captions .

长视频生成帧级标注扩散模型多场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。