arXiv:2605.22051cs.CV2026-05International Conf…被引 2

通过频域解耦,用极少资源生成高质量视觉特效

EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation

论文配图:EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation
图 1 · 摘自论文原文
  • 将画面高频纹理与低频动态分离,降低学习难度
  • 仅需约100步测试时训练即可适配新特效
  • 适合资源有限的创作者快速生成专业级特效

生成高保真视觉效果通常需要海量数据和高昂算力,因为空间纹理与时间动态高度耦合。本文提出EasyVFX,一种资源高效的框架,在严格约束下实现逼真VFX合成。核心思想是频域分解:将代表复杂空间外观的高频成分与包含全局运动动态的低频成分解耦,使高维学习问题变为可管理的子任务,降低优化门槛并减少对数据的依赖。基于此,我们设计两阶段训练范式:首先构建频率感知混合专家(Freq-MoE)架构,通过软路由机制将专用专家分配至不同频段,分别学习外观与运动先验,从而以较少GPU资源获取基础VFX知识;其次引入测试时训练策略,结合新型频域约束损失,使预训练模型可在单个GPU上仅用约100步快速适应未见特效。实验表明,EasyVFX生成的特效结构一致、视觉惊艳,证明频域感知学习是推动专业级VFX普惠的关键。

原文摘要 · Abstract (English)

Generating high-fidelity visual effects (VFX) typically demands massive datasets and prohibitive computational power due to the intricate coupling of spatial textures and temporal dynamics. In this paper, we introduce EasyVFX, a resource-efficient framework that achieves realistic VFX synthesis under stringent constraints. Our core philosophy lies in frequency-domain decomposition: we observe that the complexity of VFX can be significantly mitigated by decoupling high-frequency components, which represent intricate spatial appearances, from low-frequency components that encapsulate global motion dynamics. This spectral disentanglement transforms a high-dimensional learning problem into manageable sub-tasks, thereby lowering the optimization barrier and reducing data dependency. Building upon this insight, we propose a two-stage training paradigm. First, we design a Frequency-aware Mixture-of-Experts (Freq-MoE) architecture. By utilizing a soft routing mechanism, our model assigns specialized experts to distinct spectral bands, enabling them to cultivate robust priors for appearance and motion dynamics. This specialization allows the model to acquire foundational VFX knowledge with fewer GPU resources. Second, we introduce a Test-Time Training strategy powered by a novel Frequency-constraint Loss. This allows the pre-trained model to swiftly adapt to specific, unseen effects through localized optimizations, requiring only about 100 steps on a single GPU. Experimental results demonstrate that EasyVFX produces structurally consistent and visually stunning effects, proving that frequency-aware learning is a key catalyst for democratizing professional-grade VFX.

视觉特效频域解耦高效生成测试时训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。