提出FreeMask,解决视频编辑中注意力掩码模糊问题
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
- 引入掩码匹配成本衡量掩码随时间与模型结构的波动
- 通过自适应选择最优掩码提升语义保真度与时序一致性
- 无需微调即可集成,适合零样本视频编辑场景
文本到视频扩散模型取得显著进展,其生成连贯视频的能力推动了零样本视频编辑研究的快速发展。为提升编辑质量,常采用结构化控制,其中交叉注意力掩码控制因高效而突出。然而,直接应用于视频编辑时,该方法易引入模糊和闪烁等伪影。实验发现,此前研究忽视了一个关键因素:交叉注意力掩码并非始终清晰,而是随模型结构和去噪步骤变化。为此,本文提出掩码匹配成本(MMC)量化此波动,并设计FreeMask方法,针对具体任务选择最优掩码。利用MMC筛选的掩码,进一步优化综合注意力特征中的掩码融合机制,涵盖时间、交叉与自注意力模块。该方法可无缝集成至现有零样本视频编辑框架,无需控制辅助或参数微调,实现未编辑语义布局的自适应解耦与掩码精度控制。大量实验表明,FreeMask在语义保真度、时序一致性和编辑质量上均优于当前最先进方法。
原文摘要 · Abstract (English)
Text-to-video diffusion models have made remarkable advancements. Driven by their ability to generate temporally coherent videos, research on zero-shot video editing using these fundamental models has expanded rapidly. To enhance editing quality, structural controls are frequently employed in video editing. Among these techniques, cross-attention mask control stands out for its effectiveness and efficiency. However, when cross-attention masks are naively applied to video editing, they can introduce artifacts such as blurring and flickering. Our experiments uncover a critical factor overlooked in previous video editing research: cross-attention masks are not consistently clear but vary with model structure and denoising timestep. To address this issue, we propose the metric Mask Matching Cost (MMC) that quantifies this variability and propose FreeMask, a method for selecting optimal masks tailored to specific video editing tasks. Using MMC-selected masks, we further improve the masked fusion mechanism within comprehensive attention features, e.g., temp, cross, and self-attention modules. Our approach can be seamlessly integrated into existing zero-shot video editing frameworks with better performance, requiring no control assistance or parameter fine-tuning but enabling adaptive decoupling of unedited semantic layouts with mask precision control. Extensive experiments demonstrate that FreeMask achieves superior semantic fidelity, temporal consistency, and editing quality compared to state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。