arXiv:2510.08377cs.CV2025-10被引 84

统一视频理解生成编辑,一句指令通吃多种任务。

UniVideo: Unified Understanding, Generation, and Editing for Videos

  • 双流架构:语言模型理解指令,扩散模型生成视频。
  • 跨任务联合训练,生成与编辑效果达顶尖水平。
  • 支持指令组合与未见任务迁移,适合多模态研究者。

统一多模态模型在图像生成与编辑中表现优异,但尚未扩展至视频领域。本文提出UniVideo,一种面向视频的统一框架,采用双流设计:融合多模态大语言模型(MLLM)用于指令理解,结合多模态DiT(MMDiT)实现视频生成。该设计保留MLLM原文本生成能力,精准解析复杂多模态指令,并确保生成内容视觉一致性。基于此架构,UniVideo在单一多模态指令范式下统一多种视频生成与编辑任务,进行联合训练。大量实验表明,其在文本/图像到视频生成、上下文视频生成及编辑任务上达到或超越现有专用基线。特别地,统一设计带来两种泛化能力:一是支持任务组合(如编辑+风格迁移),通过单个指令整合多项功能;二是即使未显式训练自由形式视频编辑,仍可从大规模图像编辑数据迁移编辑能力,处理如改变环境或材质等未见指令。此外,支持基于视觉提示的视频生成,其中MLLM解析视觉提示并指导MMDiT合成。为推动后续研究,模型与代码已公开。

原文摘要 · Abstract (English)

Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain. In this work, we present UniVideo, a versatile framework that extends unified modeling to the video domain. UniVideo adopts a dual-stream design, combining a Multimodal Large Language Model (MLLM) for instruction understanding with a Multimodal DiT (MMDiT) for video generation. This design preserves the MLLM's original text generation capabilities, enables accurate interpretation of complex multimodal instructions, and maintains visual consistency in the generated content. Built on this architecture, UniVideo unifies diverse video generation and editing tasks under a single multimodal instruction paradigm and is jointly trained across them. Extensive experiments demonstrate that UniVideo matches or surpasses state-of-the-art task-specific baselines in text/image-to-video generation, in-context video generation and in-context video editing. Notably, the unified design of UniVideo enables two forms of generalization. First, UniVideo supports task composition, such as combining editing with style transfer, by integrating multiple capabilities within a single instruction. Second, even without explicit training on free-form video editing, UniVideo transfers its editing capability from large-scale image editing data to this setting, handling unseen instructions such as changing the environment or altering materials within a video. Beyond these core capabilities, UniVideo also supports visual-prompt-based video generation, where the MLLM interprets visual prompts and guides the MMDiT during synthesis. To foster future research, we released our model and code.

视频生成统一建模多模态扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。