arXiv:2505.18434cs.CVcs.AI2025-05被引 3

用训练时生成否定数据,让CLIP更好理解‘不存在’

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP

  • 训练时动态生成含否定的图像描述,仅多花2.5%时间
  • 在多个否定理解任务中达到当前最佳性能
  • 首个针对文本到图像生成的否定评估基准

视觉语言模型(如CLIP)在众多下游任务中表现优异,但在理解否定概念(即识别某概念的缺失或排除)方面仍存在局限。现有方法通过大语言模型生成大量含否定的图像描述用于微调CLIP,但此类方法耗时且计算开销大,且评估仅限于图文匹配任务。为此,我们提出:(1)一种训练时否定数据生成流程,可在训练阶段生成否定描述,仅增加2.5%额外训练时间;(2)首个评估文本到图像生成模型在否定提示下生成语义准确图像的基准Neg-TtoI。实验表明,所提方法TNG-CLIP在图像到文本匹配、文本到图像检索及图像生成等多样化否定基准上均取得当前最优表现。

原文摘要 · Abstract (English)

Vision-language models (VLMs), such as CLIP, have demonstrated strong performance across a range of downstream tasks. However, CLIP is still limited in negation understanding: the ability to recognize the absence or exclusion of a concept. Existing methods address the problem by using a large language model (LLM) to generate large-scale data of image captions containing negation for further fine-tuning CLIP. However, these methods are both time- and compute-intensive, and their evaluations are typically restricted to image-text matching tasks. To expand the horizon, we (1) introduce a training-time negation data generation pipeline such that negation captions are generated during the training stage, which only increases 2.5% extra training time, and (2) we propose the first benchmark, Neg-TtoI, for evaluating text-to-image generation models on prompts containing negation, assessing model's ability to produce semantically accurate images. We show that our proposed method, TNG-CLIP, achieves SOTA performance on diverse negation benchmarks of image-to-text matching, text-to-image retrieval, and image generation.

CLIP否定理解数据生成图文生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。