通过注意力监督实现视频生成的精准条件控制,提升可控性。
ACD: Direct Conditional Control for Video Diffusion Models via Attention Supervision
- 用注意力图对齐外部控制信号,直接约束生成过程。
- 在多个基准数据集上显著提升条件对齐度,保持时间连贯性。
- 适合需要精确动作/场景控制的视频生成应用。
可控性是视频合成的基本要求,准确对齐条件信号至关重要。现有无分类器引导方法通常通过建模数据与条件的联合分布间接实现条件控制,常导致对指定条件的控制能力有限。基于分类器的引导通过外部分类器强制条件,但模型可能仅提升分类器得分而未真正满足目标条件,产生对抗性伪影且有效控制能力不足。本文提出注意力条件扩散(ACD)框架,通过注意力监督实现视频扩散模型的直接条件控制。通过将模型注意力图与外部控制信号对齐,ACD获得更优的可控性。为此,我们引入稀疏3D感知物体布局作为高效条件信号,配套设计专用布局ControlNet及自动化标注流程,支持可扩展的布局集成。在多个基准视频生成数据集上的实验表明,ACD在保持时间连贯性和视觉保真度的同时,显著提升对条件输入的对齐能力,建立了一种有效的条件视频合成范式。
原文摘要 · Abstract (English)
Controllability is a fundamental requirement in video synthesis, where accurate alignment with conditioning signals is essential. Existing classifier-free guidance methods typically achieve conditioning indirectly by modeling the joint distribution of data and conditions, which often results in limited controllability over the specified conditions. Classifier-based guidance enforces conditions through an external classifier, but the model may exploit this mechanism to raise the classifier score without genuinely satisfying the intended condition, resulting in adversarial artifacts and limited effective controllability. In this paper, we propose Attention-Conditional Diffusion (ACD), a novel framework for direct conditional control in video diffusion models via attention supervision. By aligning the model's attention maps with external control signals, ACD achieves better controllability. To support this, we introduce a sparse 3D-aware object layout as an efficient conditioning signal, along with a dedicated Layout ControlNet and an automated annotation pipeline for scalable layout integration. Extensive experiments on benchmark video generation datasets demonstrate that ACD delivers superior alignment with conditioning inputs while preserving temporal coherence and visual fidelity, establishing an effective paradigm for conditional video synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。