arXiv:2503.18150cs.CV2025-03CVPR被引 16

无需训练即可生成高质量长视频,解决时序一致性难题。

LongDiff: Training-Free Long Video Generation in One Go

论文配图:LongDiff: Training-Free Long Video Generation in One Go
图 1 · 摘自论文原文
  • 通过位置映射与关键帧选择,实现无训练长视频生成
  • 在多个数据集上生成长达60秒的视频,保持时序连贯性
  • 适合希望直接使用现成模型生成长视频的研究者

视频扩散模型在视频生成任务中取得了显著进展。然而,大多数模型主要针对短视频设计和训练,导致在生成长视频时难以维持时序一致性和视觉细节。本文提出 LongDiff,一种无需训练的新方法,包含精心设计的两个组件——位置映射(Position Mapping, PM)和信息关键帧选择(Informative Frame Selection, IFS),以应对短到长视频生成泛化中的两大挑战:时序位置模糊与信息稀释。LongDiff 充分发挥现有视频扩散模型的能力,实现一次生成高质量长视频。大量实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Video diffusion models have recently achieved remarkable results in video generation. Despite their encouraging performance, most of these models are mainly designed and trained for short video generation, leading to challenges in maintaining temporal consistency and visual details in long video generation. In this paper, we propose LongDiff, a novel training-free method consisting of carefully designed components \ -- Position Mapping (PM) and Informative Frame Selection (IFS) \ -- to tackle two key challenges that hinder short-to-long video generation generalization: temporal position ambiguity and information dilution. Our LongDiff unlocks the potential of off-the-shelf video diffusion models to achieve high-quality long video generation in one go. Extensive experiments demonstrate the efficacy of our method.

视频生成扩散模型长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。