arXiv:2506.01801cs.CV2025-06被引 19

OmniV2V实现多种视频生成与编辑操作的统一动态控制。

OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation

  • 通过动态内容注入模块统一处理多种视频任务
  • 在多类任务中表现优于或媲美顶尖开源/商用模型
  • 适合需要灵活视频编辑的创作者与开发者

扩散变换器(DiT)的兴起推动了视频生成技术的发展,尤其在文本到视频和图像到视频任务中。然而,现有模型大多局限于单一场景,难以通过动态内容操控实现多样化的视频生成与编辑。本文提出OmniV2V,一个可跨场景生成与编辑视频的模型,支持对象移动、添加、掩码引导编辑、试穿、修复、扩展、人物动画及可控角色视频合成等操作。我们设计了一种统一的动态内容操控注入模块,有效整合各类任务需求;基于LLaVA构建视觉-文本指令理解模块,增强模型对视觉内容与指令对应关系的理解;并搭建多任务数据处理系统,解决任务间数据重叠问题,实现高效数据增强。由此构建了多类型、多场景的OmniV2V数据集及其测试基准。大量实验表明,OmniV2V在多项视频生成与编辑任务中表现不逊于甚至优于现有最佳开源与商业模型。

原文摘要 · Abstract (English)

The emergence of Diffusion Transformers (DiT) has brought significant advancements to video generation, especially in text-to-video and image-to-video tasks. Although video generation is widely applied in various fields, most existing models are limited to single scenarios and cannot perform diverse video generation and editing through dynamic content manipulation. We propose OmniV2V, a video model capable of generating and editing videos across different scenarios based on various operations, including: object movement, object addition, mask-guided video edit, try-on, inpainting, outpainting, human animation, and controllable character video synthesis. We explore a unified dynamic content manipulation injection module, which effectively integrates the requirements of the above tasks. In addition, we design a visual-text instruction module based on LLaVA, enabling the model to effectively understand the correspondence between visual content and instructions. Furthermore, we build a comprehensive multi-task data processing system. Since there is data overlap among various tasks, this system can efficiently provide data augmentation. Using this system, we construct a multi-type, multi-scenario OmniV2V dataset and its corresponding OmniV2V-Test benchmark. Extensive experiments show that OmniV2V works as well as, and sometimes better than, the best existing open-source and commercial models for many video generation and editing tasks.

视频生成动态编辑扩散模型多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。