无需反演的高效图像编辑框架,实现高保真语义精准修改。
FIA-Edit: Frequency-Interactive Attention for Efficient and High-Fidelity Inversion-Free Text-Guided Image Editing

- 通过频率交互注意力融合源图与目标特征,增强跨域对齐。
- 在RTX 4090上每张512×512图像仅需约6秒,计算成本低。
- 首次应用于临床图像编辑,助力医学数据增强与出血分类。
文本引导的图像编辑随扩散模型快速发展。尽管基于流的无反演方法因省去潜在空间反演而具备高效率,但常因无法有效整合源信息,导致背景保留差、空间不一致及过度编辑。本文提出FIA-Edit,一种新颖的无反演框架,通过频率交互注意力实现高保真与语义精确编辑。具体设计两个核心组件:(1) 频率表示交互(FRI)模块,在自注意力中交换源图与目标特征的频域成分,提升跨域对齐;(2) 特征注入(FIJ)模块,将源侧的查询、键、值及文本嵌入显式注入目标分支的交叉注意力中,以保留结构与语义。大量实验表明,FIA-Edit在低计算开销下(约6秒/张512×512图像,RTX 4090)持续优于现有方法,涵盖视觉质量、背景保真度和可控性。此外,首次将文本引导图像编辑拓展至临床应用,通过合成解剖学一致的手术图像出血变化,为医学数据增强提供新可能,并显著提升下游出血分类性能。
原文摘要 · Abstract (English)
Text-guided image editing has advanced rapidly with the rise of diffusion models. While flow-based inversion-free methods offer high efficiency by avoiding latent inversion, they often fail to effectively integrate source information, leading to poor background preservation, spatial inconsistencies, and over-editing due to the lack of effective integration of source information. In this paper, we present FIA-Edit, a novel inversion-free framework that achieves high-fidelity and semantically precise edits through a Frequency-Interactive Attention. Specifically, we design two key components: (1) a Frequency Representation Interaction (FRI) module that enhances cross-domain alignment by exchanging frequency components between source and target features within self-attention, and (2) a Feature Injection (FIJ) module that explicitly incorporates source-side queries, keys, values, and text embeddings into the target branch's cross-attention to preserve structure and semantics. Comprehensive and extensive experiments demonstrate that FIA-Edit supports high-fidelity editing at low computational cost (~6s per 512 * 512 image on an RTX 4090) and consistently outperforms existing methods across diverse tasks in visual quality, background fidelity, and controllability. Furthermore, we are the first to extend text-guided image editing to clinical applications. By synthesizing anatomically coherent hemorrhage variations in surgical images, FIA-Edit opens new opportunities for medical data augmentation and delivers significant gains in downstream bleeding classification. Our project is available at: https://github.com/kk42yy/FIA-Edit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。