arXiv:2508.07981cs.CVcs.AI2025-08AAAI被引 29

首个可统一生成多种可控视觉特效的框架,支持指定位置的多特效合成。

Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation

  • 用专家LoRA混合架构减少不同特效间的干扰。
  • 通过空间提示机制实现特效在图像中的精确定位。
  • 适合影视制作中需要多特效协同控制的场景。

视觉特效(VFX)是现代影视制作中不可或缺的视觉增强手段。尽管视频生成模型为VFX生产提供了低成本解决方案,但现有方法受限于单特效的LoRA训练,无法实现多种特效的空间协同生成。这种局限性阻碍了对多特效复合控制的应用需求。为此,我们提出Omni-Effects,首个能够生成提示引导与空间可控复合特效的统一框架。其核心包含两项创新:(1) 基于LoRA的专家混合(LoRA-MoE),通过多组专家LoRA将多种特效整合进统一模型,有效缓解跨任务干扰;(2) 空间感知提示(SAP),将空间掩码信息融入文本令牌,实现精确空间控制;并引入独立信息流(IIF)模块,隔离各特效的控制信号,防止不必要融合。为支持研究,我们构建了综合性数据集Omni-VFX,采用图像编辑与首尾帧到视频(FLF2V)合成的新型采集流程,并设计专用评估框架验证模型性能。大量实验表明,Omni-Effects可实现精准空间控制与多样特效生成,用户可同时指定特效类别与位置。

原文摘要 · Abstract (English)

Visual effects (VFX) are essential visual enhancements fundamental to modern cinematic production. Although video generation models offer cost-efficient solutions for VFX production, current methods are constrained by per-effect LoRA training, which limits generation to single effects. This fundamental limitation impedes applications that require spatially controllable composite effects, i.e., the concurrent generation of multiple effects at designated locations. However, integrating diverse effects into a unified framework faces major challenges: interference from effect variations and spatial uncontrollability during multi-VFX joint training. To tackle these challenges, we propose Omni-Effects, a first unified framework capable of generating prompt-guided effects and spatially controllable composite effects. The core of our framework comprises two key innovations: (1) LoRA-based Mixture of Experts (LoRA-MoE), which employs a group of expert LoRAs, integrating diverse effects within a unified model while effectively mitigating cross-task interference. (2) Spatial-Aware Prompt (SAP) incorporates spatial mask information into the text token, enabling precise spatial control. Furthermore, we introduce an Independent-Information Flow (IIF) module integrated within the SAP, isolating the control signals corresponding to individual effects to prevent any unwanted blending. To facilitate this research, we construct a comprehensive VFX dataset Omni-VFX via a novel data collection pipeline combining image editing and First-Last Frame-to-Video (FLF2V) synthesis, and introduce a dedicated VFX evaluation framework for validating model performance. Extensive experiments demonstrate that Omni-Effects achieves precise spatial control and diverse effect generation, enabling users to specify both the category and location of desired effects.

视觉特效空间控制统一生成LoRA-MoE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。