arXiv:2507.18046cs.CVcs.AI2025-07被引 1

通过后训练提升视频生成对场景切换的感知能力

Enhancing Scene Transition Awareness in Video Generation via Post-Training

  • 构建包含多场景切换的TAV数据集,用于后训练
  • 后训练后生成视频的场景数量更贴近提示要求
  • 适合需要长视频多场景生成的研究者使用

近期人工智能视频生成在短时单场景文本到视频任务中表现优异,但当前模型难以生成具有连贯场景切换的长视频,主要因无法从提示中推断何时需要场景转换。大多数开源模型在仅含单场景视频片段的数据集上训练,限制了其对多场景提示的理解能力。为解决此问题,我们提出了 extbf{Transition-Aware Video}(TAV)数据集,包含预处理的多场景切换视频片段。实验表明,基于TAV数据集进行后训练可显著提升模型对提示中场景切换的理解能力,缩小生成场景数与提示需求之间的差距,同时保持高质量图像输出。

原文摘要 · Abstract (English)

Recent advances in AI-generated video have shown strong performance on \emph{text-to-video} tasks, particularly for short clips depicting a single scene. However, current models struggle to generate longer videos with coherent scene transitions, primarily because they cannot infer when a transition is needed from the prompt. Most open-source models are trained on datasets consisting of single-scene video clips, which limits their capacity to learn and respond to prompts requiring multiple scenes. Developing scene transition awareness is essential for multi-scene generation, as it allows models to identify and segment videos into distinct clips by accurately detecting transitions. To address this, we propose the \textbf{Transition-Aware Video} (TAV) dataset, which consists of preprocessed video clips with multiple scene transitions. Our experiment shows that post-training on the \textbf{TAV} dataset improves prompt-based scene transition understanding, narrows the gap between required and generated scenes, and maintains image quality.

视频生成场景切换后训练数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。