arXiv:2512.19661cs.CV2025-12被引 4

让视频特效自动融合背景,无需逐帧标注

Over++: Generative Video Compositing for Layer Interaction Effects

  • 基于文本提示生成半透明环境特效,保留原视频内容
  • 仅用少量数据训练,仍生成多样且逼真的光影效果
  • 支持可选掩码和关键帧控制,适合影视后期创作者

在专业视频合成流程中,艺术家需手动添加前景与背景间的环境交互效果,如阴影、反射、灰尘和飞溅等。现有视频生成模型难以在保持输入视频的同时添加这些效果,而当前视频修复方法要么需要昂贵的逐帧掩码,要么结果不自然。我们提出增强合成这一新任务:在不破坏原场景的前提下,根据文本提示生成逼真且半透明的环境效果。为此,我们构建了一个专为此任务设计的成对特效数据集,并引入无配对增强策略以保持文本可控性。Over++ 框架无需假设相机姿态、场景静止或深度监督,支持可选掩码控制和关键帧引导,且无需密集标注。尽管训练数据有限,Over++ 仍能生成多样化、真实的环境效果,在特效生成与场景保留上均优于现有基线。

原文摘要 · Abstract (English)

In professional video compositing workflows, artists must manually create environmental interactions-such as shadows, reflections, dust, and splashes-between foreground subjects and background layers. Existing video generative models struggle to preserve the input video while adding such effects, and current video inpainting methods either require costly per-frame masks or yield implausible results. We introduce augmented compositing, a new task that synthesizes realistic, semi-transparent environmental effects conditioned on text prompts and input video layers, while preserving the original scene. To address this task, we present Over++, a video effect generation framework that makes no assumptions about camera pose, scene stationarity, or depth supervision. We construct a paired effect dataset tailored for this task and introduce an unpaired augmentation strategy that preserves text-driven editability. Our method also supports optional mask control and keyframe guidance without requiring dense annotations. Despite training on limited data, Over++ produces diverse and realistic environmental effects and outperforms existing baselines in both effect generation and scene preservation.

视频合成生成模型特效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。