SAGE用结构引导生成跨片段流畅视频过渡,无需训练。
SAGE: Structure-Aware Generative Video Transitions between Diverse Clips
- 结合轮廓线图与运动流提供结构指导,生成时保持动作连贯。
- 在多样片段间过渡效果优于现有方法,用户评分更高。
- 零样本设计,适合创意剪辑,无需复杂数据准备。
视频过渡旨在生成两段视频之间的中间帧,但简单方法(如线性混合)会引入伪影,影响专业使用或破坏时间连贯性。传统技术(交叉淡入淡出、变形、帧插值)和近期生成式中间帧方法虽可生成高质量合理中间帧,但在处理存在大时间间隔或显著语义差异的多样化片段时仍表现不佳,缺乏内容感知与视觉连贯性。本文受艺术创作流程启发,提炼对齐轮廓、插值显著特征等策略,提出SAGE(Structure-Aware Generative vidEo transitions)——一种简单有效的零样本方法。该方法利用线图与运动流提供结构引导,结合生成合成,实现无需微调的平滑、运动一致过渡。大量实验及与FILM、TVG、DiffMorpher、VACE、GI等当前方法的对比表明,SAGE在定量指标和用户评估中均优于经典与最新生成基线,在多样化片段间生成过渡效果更优。该方法有效规避了获取合适训练数据的难题,尤其适用于涉及多样片段的创造性场景。代码已公开于https://kan32501.github.io/sage.github.io/。
原文摘要 · Abstract (English)
Video transitions aim to synthesize intermediate frames between two clips, but naive approaches such as linear blending introduce artifacts that limit professional use or break temporal coherence. Traditional techniques (cross-fades, morphing, frame interpolation) and recent generative inbetweening methods can produce high-quality plausible intermediates, but they struggle with bridging diverse clips involving large temporal gaps or significant semantic differences, leaving a gap for content-aware and visually coherent transitions. We address this challenge by drawing on artistic workflows, distilling strategies such as aligning silhouettes and interpolating salient features to preserve structure and perceptual continuity. Building on these strategies, we propose SAGE (Structure-Aware Generative vidEo transitions) as a simple yet effective zeroshot approach that combines structural guidance, provided via line maps and motion flow, with generative synthesis, enabling smooth, motion-consistent transitions without fine-tuning. Extensive experiments and comparison with current alternatives, namely [FILM, TVG, DiffMorpher, VACE, GI], demonstrate that SAGE outperforms both classical and the latest generative baselines on quantitative metrics and user studies for producing transitions between diverse clips. The simple method effectively bypasses the need to acquire suitable training data, which is particularly difficult in our creative setting involving diverse clips. Code is available via the project page at https://kan32501.github.io/sage.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。