arXiv:2504.10825cs.CV2025-04AAAI被引 23

一个模型搞定视频生成与理解,还能自由切换控制方式。

OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding

  • 统一用颜色空间建模多种视觉模态,动态调整其生成或控制角色。
  • 支持文本、深度图等细粒度输入生成视频,生成质量达顶尖水平。
  • 适合视频翻译、场景重建等下游任务,灵活性强易扩展。

本文提出一种新型可控视频扩散框架 OmniVDiff,旨在通过单一扩散模型合成与理解多种视频视觉内容。OmniVDiff 将所有视频视觉模态统一在颜色空间中学习联合分布,并采用自适应控制策略,在扩散过程中动态调整各模态的角色(生成或条件)。该框架具备三项核心能力:(1) 文本引导视频生成,从文本提示联合生成所有模态;(2) 视频理解,从 RGB 输入一致预测结构化模态;(3) 多种细粒度条件下的视频生成,如深度图、Canny 边缘和分割图。大量实验表明,OmniVDiff 在视频生成任务上达到最先进性能,在视频理解任务上表现竞争力。其灵活性与可扩展性使其适用于视频到视频转换、视觉任务的模态适配及场景重建等下游应用。

原文摘要 · Abstract (English)

In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats all video visual modalities in the color space to learn a joint distribution, while employing an adaptive control strategy that dynamically adjusts the role of each visual modality during the diffusion process, either as a generation modality or a conditioning modality. Our framework supports three key capabilities: (1) Text-conditioned video generation, where all modalities are jointly synthesized from a textual prompt; (2) Video understanding, where structural modalities are predicted from rgb inputs in a coherent manner; and (3) X-conditioned video generation, where video synthesis is guided by finegrained inputs such as depth, canny and segmentation. Extensive experiments demonstrate that OmniVDiff achieves state-of-the-art performance in video generation tasks and competitive results in video understanding. Its flexibility and scalability make it well-suited for downstream applications such as video-to-video translation, modality adaptation for visual tasks, and scene reconstruction.

视频生成扩散模型可控生成多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。