arXiv:2508.07755cs.CV2025-08中稿 · CVPR

无需提示词,通过对比图像自动提取共性特征生成定制化图像

Comparison Reveals Commonality: Customized Image Generation through Contrastive Inversion

  • 通过图像间对比学习,自动识别共同语义特征
  • 在小样本图像集上实现高保真概念表示与编辑能力
  • 适合需要精准控制生成内容的个性化图像设计场景

当前对定制化图像生成的需求催生了从少量图像中有效提取共性概念的技术需求。现有方法通常依赖额外引导信息(如文本提示或空间掩码)来捕捉目标概念,但人工提供的引导可能导致辅助特征分离不完整,降低生成质量。本文提出对比反演(Contrastive Inversion)方法,通过比较输入图像本身来识别共性概念,无需额外信息。我们通过对比学习训练目标标记与图像级辅助文本标记,以提取解耦后的真正语义。随后采用解耦交叉注意力微调,提升概念保真度且避免过拟合。实验结果与分析表明,该方法在概念表征与编辑性能之间达到良好平衡,优于现有技术。

原文摘要 · Abstract (English)

The recent demand for customized image generation raises a need for techniques that effectively extract the common concept from small sets of images. Existing methods typically rely on additional guidance, such as text prompts or spatial masks, to capture the common target concept. Unfortunately, relying on manually provided guidance can lead to incomplete separation of auxiliary features, which degrades generation quality.In this paper, we propose Contrastive Inversion, a novel approach that identifies the common concept by comparing the input images without relying on additional information. We train the target token along with the image-wise auxiliary text tokens via contrastive learning, which extracts the well-disentangled true semantics of the target. Then we apply disentangled cross-attention fine-tuning to improve concept fidelity without overfitting. Experimental results and analysis demonstrate that our method achieves a balanced, high-level performance in both concept representation and editing, outperforming existing techniques.

图像生成对比学习小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。