arXiv:2603.24458cs.CV2026-03被引 21

开源统一视频生成新模型,支持自由组合与复杂意图理解。

OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning

论文配图:OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning
图 1 · 摘自论文原文
  • 构建多模态输入融合框架,实现文本、图像、视频的时序联动生成。
  • 在智能视频生成任务中达到开源模型最优性能,超越现有方案。
  • 适用于需要复杂创作意图的视频生成场景,如影视脚本可视化。

尽管商业系统如Seedance-2.0在全能力视频生成上取得显著成果,但开源模型仍明显落后。多数学术模型仍高度碎片化,现有统一视频生成尝试也难以在单一框架内无缝整合多种任务。为此,我们提出OmniWeaving,一种具备强大多模态组合与推理驱动能力的全层次视频生成模型。通过大规模预训练数据集,该模型学习在时序上绑定交错的文本、多图像和视频输入,并像智能代理一样推断复杂用户意图以生成高质量视频。此外,我们提出IntelligentVBench,首个用于评估下一代智能统一视频生成的综合性基准。大量实验表明,OmniWeaving在开源统一模型中达到最先进水平。代码与模型已公开。项目页面:https://omniweaving.github.io。

原文摘要 · Abstract (English)

While proprietary systems such as Seedance-2.0 have achieved remarkable success in omni-capable video generation, open-source alternatives significantly lag behind. Most academic models remain heavily fragmented, and the few existing efforts toward unified video generation still struggle to seamlessly integrate diverse tasks within a single framework. To bridge this gap, we propose OmniWeaving, an omni-level video generation model featuring powerful multimodal composition and reasoning-informed capabilities. By leveraging a massive-scale pretraining dataset that encompasses diverse compositional and reasoning-augmented scenarios, OmniWeaving learns to temporally bind interleaved text, multi-image, and video inputs while acting as an intelligent agent to infer complex user intentions for sophisticated video creation. Furthermore, we introduce IntelligentVBench, the first comprehensive benchmark designed to rigorously assess next-level intelligent unified video generation. Extensive experiments demonstrate that OmniWeaving achieves SoTA performance among open-source unified models. The codes and model have already been publicly available. Project Page: https://omniweaving.github.io.

视频生成多模态统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。