arXiv:2411.15034cs.CVcs.LG2024-11

无需训练即可精准编辑图像,通过智能分配注意力头实现文本引导

HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads

  • 不依赖训练,通过动态路由注意力头实现文本引导
  • 在多个基准上实现高保真度与高质量图像编辑
  • 适合需要快速部署文本编辑能力的研究者和开发者

扩散变换器(DiTs)在图像生成任务中表现出强大能力。然而,多模态DiTs(MM-DiTs)的精确文本引导图像编辑仍面临重大挑战。与可利用自/交叉注意力图进行语义编辑的UNet结构不同,MM-DiTs本身缺乏显式且一致的文本引导支持,导致编辑结果与文本之间存在语义错位。本研究揭示了不同注意力头对图像语义的敏感性,并提出HeadRouter——一种无需训练的图像编辑框架,通过自适应地将文本引导路由至MM-DiTs中的不同注意力头来编辑源图像。此外,我们设计了双标记优化模块,以优化文本/图像标记表示,实现精确的语义引导与准确的区域表达。在多个基准上的实验结果表明,HeadRouter在编辑保真度与图像质量方面表现优异。

原文摘要 · Abstract (English)

Diffusion Transformers (DiTs) have exhibited robust capabilities in image generation tasks. However, accurate text-guided image editing for multimodal DiTs (MM-DiTs) still poses a significant challenge. Unlike UNet-based structures that could utilize self/cross-attention maps for semantic editing, MM-DiTs inherently lack support for explicit and consistent incorporated text guidance, resulting in semantic misalignment between the edited results and texts. In this study, we disclose the sensitivity of different attention heads to different image semantics within MM-DiTs and introduce HeadRouter, a training-free image editing framework that edits the source image by adaptively routing the text guidance to different attention heads in MM-DiTs. Furthermore, we present a dual-token refinement module to refine text/image token representations for precise semantic guidance and accurate region expression. Experimental results on multiple benchmarks demonstrate HeadRouter's performance in terms of editing fidelity and image quality.

图像编辑注意力路由无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。