通过双阶段扩散轨迹提升图像编辑自然度与准确性
DuET: Dual Expert Trajectories for Diffusion Image Editing

- 引入双专家轨迹机制,先文本生成再回编辑模式
- 在多个模型上显著提升指令相关性与视觉质量
- 无需训练,轻量部署,适合对编辑自然度要求高的场景
现有扩散编辑方法在每一步去噪中持续依赖源图条件,可能导致编辑不充分或结果不自然,尤其当目标场景与输入差异较大时。本文提出DuET(Dual Expert Trajectories),一种无需训练的推理方法:通过在编辑(E)与文生图(T2I)之间切换,让去噪路径先向目标分布移动,再回归编辑模式,从而在保留图像结构优势的同时增强编辑效果。该方法不修改模型权重,也不增加采样成本,在多种模型和基准测试中均显著提升指令相关性、语义保真度与感知质量。固定切换策略带来可预测的源图保留损失,且该损失非本质问题。进一步提出的可选版Selective DuET,基于轻量注意力探针信号动态路由,进一步提升保真度、自然度与抗伪影能力,同时源图保留效果与基线无明显差异。
原文摘要 · Abstract (English)
Recent diffusion editors perform diverse instruction-based edits while conditioning on the source image at every denoising step. Yet persistent source-image conditioning can limit how fully an edit is executed and how natural the result appears, especially when the target scene diverges substantially from the input. We introduce DuET (Dual Expert Trajectories), a training-free inference method that temporarily relaxes source-image conditioning by transitioning through a text-to-image phase before returning to edit mode ($\mathrm{E}\to\mathrm{T2I}\to\mathrm{E}$), allowing the denoising trajectory to move toward the target distribution while retaining the structural benefits of image-conditioned editing. Without modifying model weights or increasing sampling cost, DuET consistently improves instruction relevance, semantic fidelity, and perceptual quality across diverse models and benchmarks. Fixed switching schedules obtain these gains at a modest, predictable cost in source-image preservation; we show this cost is not fundamental. A per-edit variant, Selective DuET, routes on lightweight attention-probe signals read from the edit trajectory and improves fidelity, naturalness, and artifact scores while keeping source preservation perceptually indistinguishable from the baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。