arXiv:2502.03621cs.CV2025-02SIGGRAPH被引 9

用文字指令给真实视频添加动态内容,自动融合场景细节。

DynVFX: Augmenting Real Videos with Dynamic Content

  • 通过注意力机制操控生成内容位置与运动,实现自然融合。
  • 无需训练,仅需文本指令即可生成与原视频互动的新动态元素。
  • 适合影视剪辑、创意设计等需要快速添加动态特效的场景。

我们提出一种方法,可为真实世界视频添加新生成的动态内容。给定一段输入视频和用户简单的文本指令,该方法能合成与现有场景随时间自然交互的动态物体或复杂场景效果。新内容的位置、外观和运动均无缝融入原始画面,同时考虑相机运动、遮挡及与其他动态物体的交互,生成连贯且逼真的输出视频。该方法采用零样本、无需训练的框架,利用预训练的文本到视频扩散变换器生成新内容,并通过预训练的视觉-语言模型细致构想增强后的场景。具体而言,我们提出一种新颖的基于推理的方法,操控注意力机制中的特征,实现新内容的精准定位与无缝融合,同时保持原场景完整性。本方法完全自动化,仅需简单用户指令。我们在多种真实视频上展示了其有效性,涵盖涉及相机运动与物体运动的多样化对象和场景编辑。

原文摘要 · Abstract (English)

We present a method for augmenting real-world videos with newly generated dynamic content. Given an input video and a simple user-provided text instruction describing the desired content, our method synthesizes dynamic objects or complex scene effects that naturally interact with the existing scene over time. The position, appearance, and motion of the new content are seamlessly integrated into the original footage while accounting for camera motion, occlusions, and interactions with other dynamic objects in the scene, resulting in a cohesive and realistic output video. We achieve this via a zero-shot, training-free framework that harnesses a pre-trained text-to-video diffusion transformer to synthesize the new content and a pre-trained vision-language model to envision the augmented scene in detail. Specifically, we introduce a novel inference-based method that manipulates features within the attention mechanism, enabling accurate localization and seamless integration of the new content while preserving the integrity of the original scene. Our method is fully automated, requiring only a simple user instruction. We demonstrate its effectiveness on a wide range of edits applied to real-world videos, encompassing diverse objects and scenarios involving both camera and object motion.

视频生成文本生成动态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。