arXiv:2412.10783cs.CV2024-12被引 10

无需修改模型即可实现视频扩散模型的上下文学习,生成超30秒连贯多场景视频。

Video Diffusion Transformers are In-Context Learners

  • 通过拼接视频并联合标注,实现上下文生成
  • 零额外计算开销下生成超过30秒一致视频
  • 适合需要高效可控视频生成的研究与产品开发

本文研究如何在极少调优条件下激活视频扩散模型的上下文学习能力。提出简单流程:(i)沿空间或时间维度拼接视频;(ii)从单一来源联合生成多场景视频字幕;(iii)使用精心构建的小规模数据集进行任务特定微调。通过多种可控任务验证,现有先进文本到视频模型可有效实现上下文生成。显著成果包括生成超过30秒的连贯多场景视频,且无需额外计算开销。该方法无需修改原模型,输出视频保真度高,更贴合提示要求并维持角色一致性。框架对研究社区具重要价值,为产品级可控视频生成提供关键洞见。数据、代码与模型权重已公开于:https://github.com/feizc/Video-In-Context。

原文摘要 · Abstract (English)

This paper investigates a solution for enabling in-context capabilities of video diffusion transformers, with minimal tuning required for activation. Specifically, we propose a simple pipeline to leverage in-context generation: ($\textbf{i}$) concatenate videos along spacial or time dimension, ($\textbf{ii}$) jointly caption multi-scene video clips from one source, and ($\textbf{iii}$) apply task-specific fine-tuning using carefully curated small datasets. Through a series of diverse controllable tasks, we demonstrate qualitatively that existing advanced text-to-video models can effectively perform in-context generation. Notably, it allows for the creation of consistent multi-scene videos exceeding 30 seconds in duration, without additional computational overhead. Importantly, this method requires no modifications to the original models, results in high-fidelity video outputs that better align with prompt specifications and maintain role consistency. Our framework presents a valuable tool for the research community and offers critical insights for advancing product-level controllable video generation systems. The data, code, and model weights are publicly available at: https://github.com/feizc/Video-In-Context.

视频生成扩散模型上下文学习多场景视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。