通过频域引导增强早期生成信号,实现无反演图像编辑的全局修改。
Wavelet-Guided Semantic Signal Compensation for Inversion-Free Image Editing

- 利用小波分析识别并强化早期生成中的语义信号
- 在不牺牲背景一致性的前提下,提升全局属性修改能力
- 适合需要强语义改变但保留结构的图像编辑场景
文本引导图像编辑旨在根据目标提示修改视觉内容同时保持背景不变。近期无反演框架如FlowEdit在无需反演的情况下展现出强大编辑能力。然而我们发现,在某些全局属性变化下,生成轨迹在早期时间步可能无法有效偏离源分布。分析表明,在高噪声阶段,主导数据流形的趋向性会削弱文本条件方向的影响,导致全局修改有限而背景仅部分保留。受此启发,我们提出一种无反演、频率感知的语义补偿策略,在生成初期增强有效信号,同时维持背景结构一致性。该方法在不损失背景保真度的前提下提升了全局编辑能力。
原文摘要 · Abstract (English)
Text-guided image editing aims to modify visual content according to a target prompt while preserving the background. Recent inversion-free image editing frameworks such as FlowEdit have demonstrated strong editing capability without requiring inversion. Empirically, FlowEdit can achieve substantial semantic changes under appropriate hyperparameter settings. However, we observe that under certain global attribute shifts, the editing trajectory may not effectively move away from the source distribution in the early timesteps. Our analysis suggests that in the high-noise regime, the dominant manifold-seeking flow toward the data manifold can reduce the influence of the text-conditioned direction, leading to limited global modification while background structures remain only moderately preserved. Inspired by this observation, we propose an inversion-free, frequency-aware semantic compensation strategy that strengthens the effective signal in the early stage of generation, while maintaining structural consistency in the background. The proposed method improves global editing capacity without sacrificing background fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。