CLIP理解数量存在偏差,影响图像生成准确性。
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
- 通过文本、图像和跨模态实验,系统测试CLIP对数量的理解能力。
- 发现CLIP嵌入空间中存在显著数量偏差,输出物体数常偏离要求。
- 适合关注视觉语言模型可靠性与生成精度的研究者参考。
CLIP在图像编辑、生成、视觉问答和视频理解等下游任务中表现出强大泛化能力,但其应用常因对用户意图中物体数量的理解偏差,导致生成结果与实际需求不符。本文通过精心设计的实验设置与数据集,从文本、图像及跨模态三个角度,全面评估了CLIP对数量信息的理解能力。实验结果揭示,CLIP的嵌入空间中存在显著的数量偏差,这一现象严重影响了下游任务的可靠性。
原文摘要 · Abstract (English)
CLIP has demonstrated great versatility in adapting to various downstream tasks, such as image editing and generation, visual question answering, and video understanding. However, CLIP-based applications often suffer from misunderstandings regarding user intent, leading to discrepancies between the required number of objects and the actual outputs in image generation tasks. In this work, we empirically investigate the quantity bias in CLIP. By carefully designing different experimental settings and datasets, we comprehensively evaluate CLIP's understanding of quantity from text, image, and cross-modal perspectives. Our experimental results reveal a quantity bias in CLIP embeddings, impacting the reliability of downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。