动态调控视频生成中背景与前景的注意力,提升一致性与细节质量。
When to Lock Attention: Training-Free KV Control in Video Diffusion
- 通过检测扩散过程中的幻觉风险,动态调节背景和前景的注意力融合比例。
- 在保持背景一致性的前提下,显著提升前景细节质量,避免图像伪影。
- 无需训练、即插即用,适用于各类基于DiT的视频生成模型。
在视频编辑中,保持背景一致性同时提升前景质量仍是核心挑战。注入全图信息常导致背景伪影,而固定背景锁定则严重限制前景生成能力。为此,我们提出KV-Lock,一种专为DiT-based视频扩散模型设计的训练自由框架。核心洞察是:去噪预测的方差(幻觉度量)直接反映生成多样性,且与无分类器引导(CFG)尺度密切相关。基于此,KV-Lock利用扩散幻觉检测,动态调度两个关键组件:缓存背景键值(KVs)与新生成KVs的融合比例,以及CFG尺度。当检测到幻觉风险时,系统强化背景KV锁定并同步增强条件引导以提升前景生成,从而缓解伪影并提高生成保真度。作为无需训练、即插即用的模块,可轻松集成至任意预训练的DiT-based模型中。大量实验验证,该方法在多种视频编辑任务中均优于现有方法,实现更高的前景质量与背景保真度。
原文摘要 · Abstract (English)
Maintaining background consistency while enhancing foreground quality remains a core challenge in video editing. Injecting full-image information often leads to background artifacts, whereas rigid background locking severely constrains the model's capacity for foreground generation. To address this issue, we propose KV-Lock, a training-free framework tailored for DiT-based video diffusion models. Our core insight is that the hallucination metric (variance of denoising prediction) directly quantifies generation diversity, which is inherently linked to the classifier-free guidance (CFG) scale. Building upon this, KV-Lock leverages diffusion hallucination detection to dynamically schedule two key components: the fusion ratio between cached background key-values (KVs) and newly generated KVs, and the CFG scale. When hallucination risk is detected, KV-Lock strengthens background KV locking and simultaneously amplifies conditional guidance for foreground generation, thereby mitigating artifacts and improving generation fidelity. As a training-free, plug-and-play module, KV-Lock can be easily integrated into any pre-trained DiT-based models. Extensive experiments validate that our method outperforms existing approaches in improved foreground quality with high background fidelity across various video editing tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。