无需训练即可自由编辑视频动作与动态交互,支持复杂场景修改。
Versatile Editing of Video Content, Actions, and Dynamics without Training
- 基于预训练文生视频模型,不改动模型内部实现
- 解决自由编辑时的低频错位与高频抖动问题
- 适合需要灵活编辑视频内容的研究者与创作者
近年来,可控视频生成取得显著进展,但对动作和动态事件的编辑,或插入影响其他物体行为的内容,仍是重大挑战。现有训练模型因难以收集相关数据而难以处理复杂编辑;现有无训练方法则仅限于结构与运动保持型编辑,无法修改运动或交互关系。本文提出 DynaEdit,一种基于预训练文生视频流模型的无训练编辑方法。该方法采用无需反演的机制,不干预模型内部,具备模型无关性。我们发现,直接套用该方法进行无约束编辑会导致严重的低频错位与高频抖动,揭示其成因并提出新机制克服。大量实验表明,DynaEdit 在复杂文本驱动视频编辑任务上达到当前最优性能,包括动作修改、引入与场景互动的物体及全局效应。
原文摘要 · Abstract (English)
Controlled video generation has seen drastic improvements in recent years. However, editing actions and dynamic events, or inserting contents that should affect the behaviors of other objects in real-world videos, remains a major challenge. Existing trained models struggle with complex edits, likely due to the difficulty of collecting relevant training data. Similarly, existing training-free methods are inherently restricted to structure- and motion-preserving edits and do not support modification of motion or interactions. Here, we introduce DynaEdit, a training-free editing method that unlocks versatile video editing capabilities with pretrained text-to-video flow models. Our method relies on the recently introduced inversion-free approach, which does not intervene in the model internals, and is thus model-agnostic. We show that naively attempting to adapt this approach to general unconstrained editing results in severe low-frequency misalignment and high-frequency jitter. We explain the sources for these phenomena and introduce novel mechanisms for overcoming them. Through extensive experiments, we show that DynaEdit achieves state-of-the-art results on complex text-based video editing tasks, including modifying actions, inserting objects that interact with the scene, and introducing global effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。