arXiv:2604.04487cs.CV2026-04

无需训练和反演,用视觉上下文精准编辑图像。

Training-Free Image Editing with Visual Context Integration and Concept Alignment

  • 直接基于视觉上下文转换源图,跳过繁琐反演步骤。
  • 通过概念对齐引导后验采样,显著提升编辑一致性。
  • 性能超越顶尖训练型模型,适合快速高效图像编辑。

在图像编辑中,引入上下文图像可准确传达用户需求,如主体外观或图像风格。现有基于训练的上下文感知编辑方法需大量数据收集与训练成本;而训练自由方案通常依赖扩散反演,存在一致性和灵活性不足的问题。本文提出VicoEdit,一种无需训练且无需反演的方法,将视觉上下文注入预训练的文本提示编辑模型。具体而言,VicoEdit直接根据视觉上下文将源图像转换为目标图像,避免了可能导致轨迹偏移的反演过程。此外,设计了基于概念对齐的后验采样策略以增强编辑一致性。实验证明,该训练自由方法在编辑性能上甚至优于当前最先进的训练型模型。

原文摘要 · Abstract (English)

In image editing, it is essential to incorporate a context image to convey the user's precise requirements, such as subject appearance or image style. Existing training-based visual context-aware editing methods incur data collection effort and training cost. On the other hand, the training-free alternatives are typically established on diffusion inversion, which struggles with consistency and flexibility. In this work, we propose VicoEdit, a training-free and inversion-free method to inject the visual context into the pretrained text-prompted editing model. More specifically, VicoEdit directly transforms the source image into the target one based on the visual context, thereby eliminating the need for inversion that can lead to deviated trajectories. Moreover, we design a posterior sampling approach guided by concept alignment to enhance the editing consistency. Empirical results demonstrate that our training-free method achieves even better editing performance than the state-of-the-art training-based models.

图像编辑扩散模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。