arXiv:2605.07455cs.CV2026-05被引 1

让图像编辑更忠实高效,仅靠视觉提示就能精准还原修改效果。

EditTransfer++: Toward Faithful and Efficient Visual-Prompt-Guided Image Editing

论文配图:EditTransfer++: Toward Faithful and Efficient Visual-Prompt-Guided Image Editing
图 1 · 摘自论文原文
  • 用视觉证据主导训练,摆脱文本干扰,提升编辑准确性。
  • 通过对比优化去噪路径,显著减少不一致生成,跨随机种子表现稳定。
  • 压缩条件信息复用,支持1024像素高清编辑,速度远超以往方法。

视觉提示引导的编辑迁移旨在直接从示例对中学习图像变换,相比纯文本驱动方法更具精确性和可控性。然而,现有基于扩散变压器的方法常因任务与主干结构不匹配而无法忠实复现演示编辑,问题包括预训练模型对文本条件的偏好及采样过程中的固有随机不稳定性。为此,我们提出EditTransfer++,结合渐进式结构化训练与高效条件机制,提升视觉提示的忠实度与推理效率。首先采用文本解耦训练策略,在微调阶段移除文本条件,迫使模型仅依赖视觉证据推断变换,同时在推理时仍可选择性加入文本引导;在此视觉基础模型上,引入最佳-最差对比精炼机制,重塑去噪轨迹以抑制不忠实生成,增强不同随机种子间的一致性。为缓解高分辨率上下文编辑的计算瓶颈,进一步提出条件压缩与复用策略,减少标记冗余,实现边长1024像素图像的高效生成。在现有基准和新提出的EditTransfer-Bench上的大量实验表明,EditTransfer++在视觉提示忠实度上达到当前最优水平,且推理速度显著快于先前方法,为可扩展的提示引导图像编辑及更广泛的视觉上下文学习提供了可行方向。

原文摘要 · Abstract (English)

Visual-prompt-guided edit transfer aims to learn image transformations directly from example pairs, offering more precise and controllable editing than purely text-driven approaches. However, existing diffusion transformer-based methods often fail to faithfully reproduce the demonstrated edits due to structural mismatches between the task and the backbone, including a pretrained bias toward textual conditioning and inherent stochastic instability during sampling. To bridge this gap, we present EditTransfer++, a framework that combines progressively structured training with an efficient conditioning scheme to improve both visual prompt faithfulness and inference efficiency. We first mitigate textual dominance with a text-decoupled training strategy that removes text conditioning during fine-tuning, compelling the model to infer transformations solely from visual evidence while still supporting optional text guidance at inference. On top of this visually grounded model, a best-worst contrastive refinement mechanism reshapes the denoising trajectories to suppress unfaithful generations and improve consistency across random seeds. To alleviate the computational bottleneck of high-resolution in-context editing, we further introduce a condition compression and reuse strategy that reduces token redundancy and enables efficient generation of images with a 1024-pixel long edge. Extensive experiments on existing benchmarks and the proposed EditTransfer-Bench show that EditTransfer++ achieves state-of-the-art visual prompt faithfulness with substantially faster inference than prior methods, suggesting a promising direction for scalable prompt-guided image editing and broader visual in-context learning.

图像编辑视觉提示扩散模型高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。