arXiv:2602.13637cs.CV2026-02

分治式扩散模型提升视频生成一致性

DCDM: Divide-and-Conquer Diffusion Models for Consistency-Preserving Video Generation

  • 分三模块解决片段内、片段间相机、跨镜头一致性问题
  • 在CVM竞赛测试集上显著改善视觉与语义连贯性
  • 适合需要高质量长视频生成的研究者使用

近期视频生成模型虽具备出色视觉保真度,但在语义、几何和身份一致性方面仍存挑战。本文提出系统级框架DCDM(分治扩散模型),针对三类关键问题:(1)片段内世界知识一致性,(2)片段间相机一致性,(3)跨镜头元素一致性。DCDM将一致性建模分解为三个专用组件,共享统一的视频生成主干。针对片段内一致性,利用大语言模型解析输入提示为结构化语义表示,再由扩散Transformer生成连贯内容;针对片段间相机一致性,提出噪声空间中的时序相机表征,实现精确稳定的运动控制,并引入文本到图像初始化机制提升可控性;针对跨镜头一致性,采用全局场景生成范式,结合窗口化交叉注意力与稀疏跨镜头自注意力,保障长程叙事连贯性同时保持计算效率。我们在AAAI'26 CVM竞赛测试集上验证该框架,结果表明所提策略有效解决了上述挑战。

原文摘要 · Abstract (English)

Recent video generative models have demonstrated impressive visual fidelity, yet they often struggle with semantic, geometric, and identity consistency. In this paper, we propose a system-level framework, termed the Divide-and-Conquer Diffusion Model (DCDM), to address three key challenges: (1) intra-clip world knowledge consistency, (2) inter-clip camera consistency, and (3) inter-shot element consistency. DCDM decomposes video consistency modeling under these scenarios into three dedicated components while sharing a unified video generation backbone. For intra-clip consistency, DCDM leverages a large language model to parse input prompts into structured semantic representations, which are subsequently translated into coherent video content by a diffusion transformer. For inter-clip camera consistency, we propose a temporal camera representation in the noise space that enables precise and stable camera motion control, along with a text-to-image initialization mechanism to further enhance controllability. For inter-shot consistency, DCDM adopts a holistic scene generation paradigm with windowed cross-attention and sparse inter-shot self-attention, ensuring long-range narrative coherence while maintaining computational efficiency. We validate our framework on the test set of the CVM Competition at AAAI'26, and the results demonstrate that the proposed strategies effectively address these challenges.

视频生成扩散模型一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。