无需训练,用文本生成电影级匹配剪辑。
MatchDiffusion: Training-free Generation of Match-cuts
- 共享噪声初始化结构,分阶段控制视频一致性与差异性。
- 生成视频在视觉连贯性上优于基线方法,支持自然过渡。
- 适合影视创作者快速实现创意剪辑,无需专业技能。
匹配剪辑是强大的电影手法,能实现场景间的无缝转换,带来强烈的视觉与隐喻关联。然而,创作匹配剪辑耗时耗力,需精心的艺术规划。本文提出首个无需训练的匹配剪辑生成方法——MatchDiffusion,基于文本到视频扩散模型。该方法利用扩散模型特性:早期去噪步骤决定画面整体结构,后期添加细节。MatchDiffusion采用“联合扩散”从共享噪声初始化两个提示的生成,对齐结构与运动;随后使用“分离扩散”,使视频逐渐发散并引入独特细节。该流程生成的视频具备视觉连贯性,适用于匹配剪辑。用户研究与量化指标均验证其有效性,展现了降低匹配剪辑创作门槛的潜力。
原文摘要 · Abstract (English)
Match-cuts are powerful cinematic tools that create seamless transitions between scenes, delivering strong visual and metaphorical connections. However, crafting match-cuts is a challenging, resource-intensive process requiring deliberate artistic planning. In MatchDiffusion, we present the first training-free method for match-cut generation using text-to-video diffusion models. MatchDiffusion leverages a key property of diffusion models: early denoising steps define the scene's broad structure, while later steps add details. Guided by this insight, MatchDiffusion employs "Joint Diffusion" to initialize generation for two prompts from shared noise, aligning structure and motion. It then applies "Disjoint Diffusion", allowing the videos to diverge and introduce unique details. This approach produces visually coherent videos suited for match-cuts. User studies and metrics demonstrate MatchDiffusion's effectiveness and potential to democratize match-cut creation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。