用参考视频让AI一键生成动态视觉特效,还能跨类别泛化。
VFXMaster: Unlocking Dynamic Visual Effect Generation via In-Context Learning
- 将特效生成转为上下文学习,用参考视频引导模型模仿
- 单次输入参考视频即可快速适配新特效,泛化能力显著提升
- 首个统一框架,支持多样动态特效且无信息泄露
视觉特效(VFX)对数字媒体的表现力至关重要,但其生成仍是生成式AI的重大挑战。现有方法多采用「每种特效一个LoRA」的范式,资源消耗大且无法泛化到未见特效,限制了可扩展性与创作力。为此,我们提出VFXMaster,首个基于参考的统一视觉特效视频生成框架。它将特效生成重构为上下文学习任务,仅需一个参考视频即可在目标内容上复现多种动态特效。通过设计上下文条件策略与注意力掩码,精确解耦并注入关键特效属性,实现单一模型掌握多种特效而无信息泄露。此外,提出高效的一次性特效适配机制,仅需用户提供的单个视频即可快速提升对复杂未见特效的泛化能力。大量实验表明,该方法能有效模仿多种特效类别,并展现出卓越的域外泛化性能。为推动后续研究,我们将开源代码、模型及完整数据集。
原文摘要 · Abstract (English)
Visual effects (VFX) are crucial to the expressive power of digital media, yet their creation remains a major challenge for generative AI. Prevailing methods often rely on the one-LoRA-per-effect paradigm, which is resource-intensive and fundamentally incapable of generalizing to unseen effects, thus limiting scalability and creation. To address this challenge, we introduce VFXMaster, the first unified, reference-based framework for VFX video generation. It recasts effect generation as an in-context learning task, enabling it to reproduce diverse dynamic effects from a reference video onto target content. In addition, it demonstrates remarkable generalization to unseen effect categories. Specifically, we design an in-context conditioning strategy that prompts the model with a reference example. An in-context attention mask is designed to precisely decouple and inject the essential effect attributes, allowing a single unified model to master the effect imitation without information leakage. In addition, we propose an efficient one-shot effect adaptation mechanism to boost generalization capability on tough unseen effects from a single user-provided video rapidly. Extensive experiments demonstrate that our method effectively imitates various categories of effect information and exhibits outstanding generalization to out-of-domain effects. To foster future research, we will release our code, models, and a comprehensive dataset to the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。