arXiv:2412.19978cs.CV2024-12被引 2

无需微调,用掩码调控注意力实现多属性精准视频编辑

MAKIMA: Tuning-free Multi-Attribute Open-domain Video Editing via Mask-Guided Attention Modulation

  • 通过反演特征与掩码引导注意力调节,精准控制多属性变化
  • 在多个数据集上实现更高编辑准确率与时间一致性,优于现有方法
  • 适合需要高效多属性视频编辑的创作者与研究人员

基于扩散的文本到图像模型在全局视频编辑任务中表现优异,但对属性特定修改仍具挑战,尤其在开放域多属性视频编辑(MAE)方面。现有方法或需大量微调,或依赖额外网络(如ControlNet),且仅能实现粗粒度编辑。本文提出MAKIMA,一种基于预训练T2I模型的免微调多属性开放域视频编辑框架。通过在去噪过程中引入反演得到的注意力图与特征,保留视频结构与外观信息。为实现多属性精准编辑,提出掩码引导注意力调制机制,增强空间对应标记间的关联性,抑制自注意力与交叉注意力层中的跨属性干扰。为平衡生成质量与效率,采用一致特征传播策略,仅编辑关键帧并传播特征至全序列。大量实验表明,MAKIMA在开放域多属性视频编辑任务中优于现有基线,在编辑准确率与时间一致性方面均表现更优,同时保持计算高效。

原文摘要 · Abstract (English)

Diffusion-based text-to-image (T2I) models have demonstrated remarkable results in global video editing tasks. However, their focus is primarily on global video modifications, and achieving desired attribute-specific changes remains a challenging task, specifically in multi-attribute editing (MAE) in video. Contemporary video editing approaches either require extensive fine-tuning or rely on additional networks (such as ControlNet) for modeling multi-object appearances, yet they remain in their infancy, offering only coarse-grained MAE solutions. In this paper, we present MAKIMA, a tuning-free MAE framework built upon pretrained T2I models for open-domain video editing. Our approach preserves video structure and appearance information by incorporating attention maps and features from the inversion process during denoising. To facilitate precise editing of multiple attributes, we introduce mask-guided attention modulation, enhancing correlations between spatially corresponding tokens and suppressing cross-attribute interference in both self-attention and cross-attention layers. To balance video frame generation quality and efficiency, we implement consistent feature propagation, which generates frame sequences by editing keyframes and propagating their features throughout the sequence. Extensive experiments demonstrate that MAKIMA outperforms existing baselines in open-domain multi-attribute video editing tasks, achieving superior results in both editing accuracy and temporal consistency while maintaining computational efficiency.

视频编辑扩散模型多属性免微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。