arXiv:2608.03059cs.CV2026-08

无需反演的图像编辑新方法,通过动态引导提升生成质量

RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editing

论文配图:RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editing
图 1 · 摘自论文原文
  • 用内部动态掩码在高噪声阶段引导编辑状态演化
  • 在高噪声水平下自然减小源与目标间的噪声位移
  • 不依赖外部模型,适用于主流文生图架构

无反演的基于流的图像编辑避免了潜在空间反演,但仍需每步维护目标侧状态。现有等位移构造在不同噪声水平下保持位移不变,与去噪过程不一致——相同噪声水平和噪声样本下,两个干净状态的位移应随噪声增加而收缩。这会导致高噪声水平时更新过激。我们提出RIDGE:一种无反演、免训练的图像编辑方法,将编辑状态作为不可见干净目标状态的动态近似。通过使用与原始源状态相同的噪声水平和噪声样本进行重去噪,使两者噪声位移随噪声升高自然减小。由于编辑状态初始目标语义有限,RIDGE在早期高噪声步骤中引入内部动态引导:由模型内部推导出的软动态掩码,指导预测的干净目标状态,聚焦于需修改区域,无需外部分割或检测模型。在两个基准测试上,使用SD3 Medium和FLUX.1-dev两种骨干网络的实验表明,RIDGE在源保留、目标对齐和感知质量之间实现了优良权衡。

原文摘要 · Abstract (English)

Inversion-free flow-based image editing avoids latent inversion, but still requires a target-side state at every editing step. The widely used equal-displacement construction keeps the displacement between the noisy source state and the target-side state unchanged across noise levels. This is inconsistent with noising, under which the displacement between two clean states noised with the same noise level and noise sample should contract as the noise level increases. Thus, it can lead to overly aggressive updates at high noise levels. We introduce RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editing, an inversion-free and training-free method that maintains the edited state as an evolving approximation to the unavailable clean target state. RIDGE re-noises this approximation using the same noise level and noise sample as the clean source state, allowing their noisy displacement to decrease naturally with increasing noise. Since the edited state initially contains limited target semantics, RIDGE further applies internal dynamic guidance during the early high-noise steps. A clean target state prediction guides the provisional edited state through a soft dynamic mask derived internally from the model, focusing guidance on regions that require modification without external segmentation or detection models. Experiments on two benchmarks using two backbones, SD3 Medium and FLUX.1-dev, show that RIDGE offers a favorable aggregate trade-off among source preservation, target alignment, and perceptual quality.

图像编辑扩散模型动态引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。