arXiv:2605.21190cs.CV2026-05中稿 · ICML被引 1

让图像编辑同时保持语义精准和结构稳定,不需重新训练模型。

Semantic Granularity Navigation in Image Editing

论文配图:Semantic Granularity Navigation in Image Editing
图 1 · 摘自论文原文
  • 通过控制中间尺度分配,解耦编辑进度与模型噪声强度。
  • 在多个编辑器和流模型上实现平均性能提升,且无需重训练。
  • 适用于希望提升编辑质量但无法修改模型的研究者或开发者。

尽管扩散模型和流模型具备强大的生成能力,真实图像编辑仍受限于语义可编辑性与结构保真度之间的权衡。我们发现这一限制的主要原因在于现有方法中编辑进度与模型尺度的隐式耦合:更强的编辑通常需进入更嘈杂的状态,导致计算资源浪费在破坏布局上,而非精准定位语义变化。为此,我们提出 NaviEdit,一种无需训练的推理时控制器,通过严格的自洽约束解耦编辑进度与模型尺度遍历。NaviEdit 在轨迹层面操作,不改变预训练模型,将尺度作为控制输入,将固定的步数预算重新分配至对语义更敏感的中间尺度,而非破坏性的高噪声区域。实验表明,在兼容的编辑器和流模型骨干网络上均取得正向平均增益,验证了该解耦策略作为可移植推理时控制原则的有效性。

原文摘要 · Abstract (English)

Despite the generative capabilities of diffusion and flow models, real-image editing remains constrained by a persistent trade-off between semantic editability and structural fidelity. We trace a primary cause of this limitation to the implicit coupling of edit progress with model scale in existing paradigms. Under this coupling, stronger edits typically require visiting noisier states, which spends computation on destabilizing layout before the semantic change is well localized. We introduce NaviEdit, a training-free inference-time controller that decouples edit progress from model scale traversal through a strict self-consistency contract. NaviEdit operates at the rollout level and leaves the underlying pretrained model unchanged. It treats scale as a control input and reallocates a fixed step budget toward semantically responsive intermediate scales instead of destructive high-noise regimes. Experiments show positive average gains across compatible editors and flow backbones, supporting decoupling as a portable inference-time control principle.

图像编辑扩散模型推理控制语义导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。