一个异常文本能骗过图文编码器,暴露其评估漏洞。
One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via Hubness

- 通过寻找高中心性文本,定位跨模态编码器的脆弱点。
- 单个伪造文本在多个图像上得分媲美甚至超过真人描述。
- 适合研究图文匹配、模型安全或评测机制的学者参考。
高维嵌入空间中的中心性问题(即某些嵌入与大量无关样本距离过近)常影响信息检索和自动评估指标的可靠性。由于文本与图像间的跨模态相似性无法通过字符串匹配直接计算,跨模态编码器将不同模态映射到共享空间以支持各类应用,而中心性嵌入的存在可能带来实际威胁。为揭示此类编码器的脆弱性,我们提出一种识别中心嵌入及其对应中心文本的方法。在MSCOCO与nocaps图像描述评估任务,以及MSCOCO和Flickr30k图像到文本检索任务上的实验表明,该方法可找到一个单一的中心文本,在众多图像上产生的相似度得分与人类参考句相当甚至更高,暴露出跨模态编码器的潜在缺陷。
原文摘要 · Abstract (English)
The hubness problem, in which hub embeddings are close to many unrelated examples, occurs often in high-dimensional embedding spaces and may pose a practical threat for purposes such as information retrieval and automatic evaluation metrics. In particular, since cross-modal similarity between text and images cannot be calculated by direct comparisons, such as string matching, cross-modal encoders that project different modalities into a shared space are helpful for various cross-modal applications, and thus, the existence of hubs may pose practical threats. To reveal the vulnerabilities of cross-modal encoders, we propose a method for identifying the hub embedding and its corresponding hub text. Experiments on image captioning evaluation in MSCOCO and nocaps along with image-to-text retrieval tasks in MSCOCO and Flickr30k showed that our method can identify a single hub text that unreasonably achieves comparable or higher similarity scores than human-written reference captions in many images, thereby revealing the vulnerabilities in cross-modal encoders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。