arXiv:2503.20484cs.CVcs.AI2025-03被引 3

无需训练,用对比学习实现精准图像翻译。

Contrastive Learning Guided Latent Diffusion Model for Image-to-Image Translation

  • 通过局部对比损失自动确定文本编辑方向。
  • 生成图像与原图嵌入间保持高保真度,结构更稳定。
  • 零样本适配,直接部署于预训练扩散模型上。

扩散模型在文本引导的图像翻译中表现出色,能生成多样且高质量的图像。然而,在文本提示设计和参考图像内容保持方面仍有提升空间。首先,目标文本提示的变化会显著影响生成图像质量,用户难以构建能完整捕捉输入图像内容的最优提示。其次,现有模型虽可对参考图像特定区域进行修改,但常导致应保持不变区域出现非预期变化。为此,我们提出 pix2pix-zeroCon,一种基于对比学习的零样本扩散方法,无需额外训练,通过局部对比损失实现精准控制。具体而言,我们根据参考图像和目标提示自动确定文本嵌入空间中的编辑方向。为确保生成图像的内容与结构精确保留,引入交叉注意力引导损失和生成图像与原始图像嵌入间的局部对比损失。该方法直接作用于预训练的文本到图像扩散模型,无需再训练。大量实验表明,本方法在图像到图像翻译任务中优于现有模型,显著提升了生成图像的保真度与可控性。

原文摘要 · Abstract (English)

The diffusion model has demonstrated superior performance in synthesizing diverse and high-quality images for text-guided image translation. However, there remains room for improvement in both the formulation of text prompts and the preservation of reference image content. First, variations in target text prompts can significantly influence the quality of the generated images, and it is often challenging for users to craft an optimal prompt that fully captures the content of the input image. Second, while existing models can introduce desired modifications to specific regions of the reference image, they frequently induce unintended alterations in areas that should remain unchanged. To address these challenges, we propose pix2pix-zeroCon, a zero-shot diffusion-based method that eliminates the need for additional training by leveraging patch-wise contrastive loss. Specifically, we automatically determine the editing direction in the text embedding space based on the reference image and target prompts. Furthermore, to ensure precise content and structural preservation in the edited image, we introduce cross-attention guiding loss and patch-wise contrastive loss between the generated and original image embeddings within a pre-trained diffusion model. Notably, our approach requires no additional training and operates directly on a pre-trained text-to-image diffusion model. Extensive experiments demonstrate that our method surpasses existing models in image-to-image translation, achieving enhanced fidelity and controllability.

图像翻译扩散模型对比学习零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。