arXiv:2512.07831cs.CV2025-12被引 11

统一多模态输入,让视频生成更懂物理世界规则

UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation

  • 用动态噪声统一不同模态训练流程,提升学习效率
  • 在130万样本数据上实现更快收敛与更强零样本泛化能力
  • 适合需要真实物理约束的视频生成研究者使用

近期视频生成模型虽具备强大合成能力,但受限于单一模态条件输入,难以实现全面的世界理解。这源于跨模态交互不足及模态多样性有限。为此,我们提出UnityVideo,一种面向世界感知视频生成的统一框架,联合学习多种模态(分割掩码、人体骨骼、DensePose、光流、深度图)和训练范式。方法包含两个核心组件:(1) 动态噪声机制,统一异构训练流程;(2) 带上下文学习器的模态切换器,通过模块化参数与情境学习实现统一处理。我们构建了一个包含130万样本的大规模统一数据集。联合优化后,UnityVideo加速收敛,并显著提升对未见数据的零样本泛化能力。实验表明,该模型生成视频质量更高、时序更一致,且更符合物理世界约束。代码与数据见:https://github.com/dvlab-research/UnityVideo

原文摘要 · Abstract (English)

Recent video generation models demonstrate impressive synthesis capabilities but remain limited by single-modality conditioning, constraining their holistic world understanding. This stems from insufficient cross-modal interaction and limited modal diversity for comprehensive world knowledge representation. To address these limitations, we introduce UnityVideo, a unified framework for world-aware video generation that jointly learns across multiple modalities (segmentation masks, human skeletons, DensePose, optical flow, and depth maps) and training paradigms. Our approach features two core components: (1) dynamic noising to unify heterogeneous training paradigms, and (2) a modality switcher with an in-context learner that enables unified processing via modular parameters and contextual learning. We contribute a large-scale unified dataset with 1.3M samples. Through joint optimization, UnityVideo accelerates convergence and significantly enhances zero-shot generalization to unseen data. We demonstrate that UnityVideo achieves superior video quality, consistency, and improved alignment with physical world constraints. Code and data can be found at: https://github.com/dvlab-research/UnityVideo

视频生成多模态世界感知扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。