通过域适应方法提升文本控制视频生成的画质与运动可控性
Interactive Video Generation via Domain Adaptation
- 引入掩码归一化缓解注意力遮蔽导致的分布偏移
- 设计时空一致性先验,解决初始噪声与条件不匹配问题
- 在保持轨迹控制精度的同时显著提升视频视觉质量
文本条件扩散模型已成为高质量视频生成的强大工具。然而,实现用户可交互的视频生成(IVG),即控制物体运动轨迹等动态元素,仍具挑战。现有无训练方法通过注意力遮蔽引导轨迹,但常导致感知质量下降。我们识别出两类关键失败模式,均源于域偏移问题,并提出受域适应启发的解决方案。首先,将感知退化归因于注意力遮蔽引发的内部协变量偏移,因预训练模型未学习处理遮蔽注意力。为此,我们提出掩码归一化,一种通过分布匹配减轻该偏移的预归一化层。其次,针对初始噪声与IVG条件间存在初始化差距的问题,引入时序内在扩散先验,在每一步去噪中强制保持时空一致性。大量定性和定量评估表明,掩码归一化与时序内在去噪能显著提升感知质量与轨迹控制能力,优于现有最先进的IVG技术。
原文摘要 · Abstract (English)
Text-conditioned diffusion models have emerged as powerful tools for high-quality video generation. However, enabling Interactive Video Generation (IVG), where users control motion elements such as object trajectory, remains challenging. Recent training-free approaches introduce attention masking to guide trajectory, but this often degrades perceptual quality. We identify two key failure modes in these methods, both of which we interpret as domain shift problems, and propose solutions inspired by domain adaptation. First, we attribute the perceptual degradation to internal covariate shift induced by attention masking, as pretrained models are not trained to handle masked attention. To address this, we propose mask normalization, a pre-normalization layer designed to mitigate this shift via distribution matching. Second, we address initialization gap, where the randomly sampled initial noise does not align with IVG conditioning, by introducing a temporal intrinsic diffusion prior that enforces spatio-temporal consistency at each denoising step. Extensive qualitative and quantitative evaluations demonstrate that mask normalization and temporal intrinsic denoising improve both perceptual quality and trajectory control over the existing state-of-the-art IVG techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。