用首尾帧和文本生成流畅自然的视频过渡,支持多种内容变换。
Versatile Transition Generation with Image-to-Video Diffusion
- 基于插值初始化保持物体一致性,应对内容突变。
- 双方向运动微调提升动作平滑度,对齐表征增强生成质量。
- 适用于概念融合与场景切换,适合视频编辑与创作研究者。
利用文本、图像、结构图或运动轨迹作为条件引导,扩散模型已在自动化高质量视频生成中取得显著进展。然而,仅给定首尾视频帧及描述性文本提示,生成平滑且合理的过渡视频仍鲜有研究。本文提出VTG(Versatile Transition video Generation)框架,可生成平滑、高保真且语义一致的视频过渡。VTG采用基于插值的初始化策略,有效保留物体身份并处理突发内容变化;同时引入双向运动微调与表征对齐正则化,分别缓解预训练图像到视频扩散模型在运动平滑性和生成保真度上的不足。为评估VTG并推动统一过渡生成研究,我们构建了TransitBench,一个涵盖概念融合与场景切换两类典型任务的综合性基准。大量实验表明,VTG在全部四项任务中均表现出卓越的过渡性能。
原文摘要 · Abstract (English)
Leveraging text, images, structure maps, or motion trajectories as conditional guidance, diffusion models have achieved great success in automated and high-quality video generation. However, generating smooth and rational transition videos given the first and last video frames as well as descriptive text prompts is far underexplored. We present VTG, a Versatile Transition video Generation framework that can generate smooth, high-fidelity, and semantically coherent video transitions. VTG introduces interpolation-based initialization that helps preserve object identity and handle abrupt content changes effectively. In addition, it incorporates dual-directional motion fine-tuning and representation alignment regularization to mitigate the limitations of pre-trained image-to-video diffusion models in motion smoothness and generation fidelity, respectively. To evaluate VTG and facilitate future studies on unified transition generation, we collected TransitBench, a comprehensive benchmark for transition generation covering two representative transition tasks: concept blending and scene transition. Extensive experiments show that VTG achieves superior transition performance consistently across all four tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。