arXiv:2507.01496cs.CV2025-07ICCV被引 16

无需训练和提示,用中间特征实现高质量图像编辑。

ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation

  • 利用中步潜在表示提取真实图像结构特征
  • 通过注意力自适应提升文本对齐与编辑效果
  • 无须掩码或源提示,适合快速精准编辑

修正流(Rectified Flow)文本到图像模型在图像质量和文本对齐上超越扩散模型,但将其用于真实图像编辑仍具挑战。本文分析多模态变换器块的中间表示,识别出三个关键特征。为从真实图像中提取这些特征并保留足够结构,我们仅将潜在表示反演至中步。随后在注入时自适应调整注意力机制,以增强可编辑性并提高与目标文本的对齐度。所提方法无需训练、不依赖用户提供的掩码,且可在无源提示条件下使用。在两个基准数据集上与九种基线方法对比,实验结果表明其性能更优,人类评估也证实用户偏好明显高于现有方法。

原文摘要 · Abstract (English)

Rectified Flow text-to-image models surpass diffusion models in image quality and text alignment, but adapting ReFlow for real-image editing remains challenging. We propose a new real-image editing method for ReFlow by analyzing the intermediate representations of multimodal transformer blocks and identifying three key features. To extract these features from real images with sufficient structural preservation, we leverage mid-step latent, which is inverted only up to the mid-step. We then adapt attention during injection to improve editability and enhance alignment to the target text. Our method is training-free, requires no user-provided mask, and can be applied even without a source prompt. Extensive experiments on two benchmarks with nine baselines demonstrate its superior performance over prior methods, further validated by human evaluations confirming a strong user preference for our approach.

图像编辑修正流文本引导无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。