让视频生成学会电影级转场,实现自然流畅的多镜头切换。
CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion Models
- 用掩码控制扩散模型,在不训练的情况下实现任意位置转场。
- 在Cine250K数据集上微调后,转场自然且符合电影剪辑风格。
- 提出新评估指标,全面验证转场控制与视频连贯性优势。
尽管视频合成技术取得显著进展,但多镜头视频生成仍处于起步阶段。即使模型规模扩大、数据集庞大,镜头切换能力依然粗糙且不稳定,多数生成视频仅限于单镜头序列。本文提出CineTrans框架,实现具有电影风格的连贯多镜头视频生成。我们构建了包含详细镜头标注的多镜头视频-文本数据集Cine250K,以深入理解电影剪辑风格。通过分析现有视频扩散模型,发现其注意力图与镜头边界存在对应关系,据此设计基于掩码的控制机制,可在无训练设置下实现任意位置的转场迁移。在该数据集上微调后,CineTrans生成的视频符合电影剪辑规范,避免了不稳定的转场或简单拼接。最后,我们提出专门的评估指标,用于衡量转场控制、时间一致性与整体质量,并通过大量实验表明,该方法在各项指标上均显著优于现有基线。
原文摘要 · Abstract (English)
Despite significant advances in video synthesis, research into multi-shot video generation remains in its infancy. Even with scaled-up models and massive datasets, the shot transition capabilities remain rudimentary and unstable, largely confining generated videos to single-shot sequences. In this work, we introduce CineTrans, a novel framework for generating coherent multi-shot videos with cinematic, film-style transitions. To facilitate insights into the film editing style, we construct a multi-shot video-text dataset Cine250K with detailed shot annotations. Furthermore, our analysis of existing video diffusion models uncovers a correspondence between attention maps in the diffusion model and shot boundaries, which we leverage to design a mask-based control mechanism that enables transitions at arbitrary positions and transfers effectively in a training-free setting. After fine-tuning on our dataset with the mask mechanism, CineTrans produces cinematic multi-shot sequences while adhering to the film editing style, avoiding unstable transitions or naive concatenations. Finally, we propose specialized evaluation metrics for transition control, temporal consistency and overall quality, and demonstrate through extensive experiments that CineTrans significantly outperforms existing baselines across all criteria.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。