arXiv:2602.08820cs.CV2026-02被引 17

用大模型理解指令,高效生成和编辑视频

Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing

  • 用多模态大模型解析用户指令,生成精准描述引导生成
  • 140亿参数模型支持高质量视频生成与复杂编辑任务
  • 轻量适配器实现高效参数复用,适合需要精准控制的视频创作

我们提出Omni-Video 2,一种可扩展且计算高效的模型,将预训练多模态大语言模型(MLLM)与视频扩散模型结合,实现统一的视频生成与编辑。核心思想是利用MLLM的理解与推理能力,生成明确的目标描述以解析用户指令。由此,理解模型中的丰富上下文表征被直接用于引导生成过程,显著提升对复杂组合指令的遵循能力。同时,设计轻量级适配器,将多模态条件令牌注入预训练文本到视频扩散模型,以参数高效方式最大限度复用其强大的生成先验。得益于这些设计,我们在精心筛选的训练数据上将Omni-Video 2扩展至140亿参数的视频扩散模型,支持高质量文本到视频生成及多种视频编辑任务,如物体移除、添加、背景更换、复杂运动编辑等。在FiVE细粒度视频编辑基准和VBench文本到视频生成基准上的评估显示,该模型在遵循复杂组合指令方面表现优异,同时在视频生成任务中达到或超过现有水平。

原文摘要 · Abstract (English)

We present Omni-Video 2, a scalable and computationally efficient model that connects pretrained multimodal large-language models (MLLMs) with video diffusion models for unified video generation and editing. Our key idea is to exploit the understanding and reasoning capabilities of MLLMs to produce explicit target captions to interpret user instructions. In this way, the rich contextual representations from the understanding model are directly used to guide the generative process, thereby improving performance on complex and compositional editing. Moreover, a lightweight adapter is developed to inject multimodal conditional tokens into pretrained text-to-video diffusion models, allowing maximum reuse of their powerful generative priors in a parameter-efficient manner. Benefiting from these designs, we scale up Omni-Video 2 to a 14B video diffusion model on meticulously curated training data with quality, supporting high quality text-to-video generation and various video editing tasks such as object removal, addition, background change, complex motion editing, \emph{etc.} We evaluate the performance of Omni-Video 2 on the FiVE benchmark for fine-grained video editing and the VBench benchmark for text-to-video generation. The results demonstrate its superior ability to follow complex compositional instructions in video editing, while also achieving competitive or superior quality in video generation tasks.

视频生成多模态扩散模型编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。