用上下文学习实现精准高效的视频特效编辑,保持背景不变且效果连贯。
IC-Effect: Precise and Efficient Video Effects Editing via In-Context Learning
- 基于DiT架构,利用源视频作上下文,实现精确背景保留和自然特效注入。
- 在15种高质量视觉风格数据上,仅需少量配对样本即可生成连贯特效。
- 支持火焰、粒子等复杂效果,适合影视创作与快速原型设计人群。
我们提出IC-Effect,一种基于DiT的指令引导式少样本视频视觉特效(VFX)编辑框架,能够合成火焰、粒子及卡通角色等复杂特效,同时严格保持空间与时间一致性。视频VFX编辑极具挑战性:注入的特效必须与背景无缝融合,背景须完全不变,且需从有限成对数据中高效学习特效模式。然而现有模型难以满足这些要求。IC-Effect将源视频作为干净上下文条件,利用DiT模型的上下文学习能力,实现精准背景保护与自然特效注入。采用两阶段训练策略——先进行通用编辑适应,再通过Effect-LoRA实现特效特定学习,确保强指令遵循与鲁棒特效建模。为提升效率,引入时空稀疏标记化,显著降低计算量的同时保持高保真。我们还发布了涵盖15种高质量视觉风格的成对VFX编辑数据集。大量实验表明,IC-Effect能实现高质量、可控且时间一致的VFX编辑,为视频创作开辟新可能。
原文摘要 · Abstract (English)
We propose \textbf{IC-Effect}, an instruction-guided, DiT-based framework for few-shot video VFX editing that synthesizes complex effects (\eg flames, particles and cartoon characters) while strictly preserving spatial and temporal consistency. Video VFX editing is highly challenging because injected effects must blend seamlessly with the background, the background must remain entirely unchanged, and effect patterns must be learned efficiently from limited paired data. However, existing video editing models fail to satisfy these requirements. IC-Effect leverages the source video as clean contextual conditions, exploiting the contextual learning capability of DiT models to achieve precise background preservation and natural effect injection. A two-stage training strategy, consisting of general editing adaptation followed by effect-specific learning via Effect-LoRA, ensures strong instruction following and robust effect modeling. To further improve efficiency, we introduce spatiotemporal sparse tokenization, enabling high fidelity with substantially reduced computation. We also release a paired VFX editing dataset spanning $15$ high-quality visual styles. Extensive experiments show that IC-Effect delivers high-quality, controllable, and temporally consistent VFX editing, opening new possibilities for video creation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。