CLIP文本编码器对微小输入改动敏感,影响检索稳定性。
On the Brittleness of CLIP Text Encoders
- 分析多种非语义文本扰动对CLIP检索的影响
- 语法和语义扰动导致最大结果波动,标点大小写也引发不稳
- 提示词微调易出错,适合关注模型鲁棒性的研究者
多模态联合嵌入模型,尤其是CLIP,近年来通过将图像与文本对齐在共享表示空间中,显著提升了零样本分类和多媒体信息检索的性能。然而,这类基于对比学习训练的模型对微小输入扰动缺乏稳定性。尤其在人工表达的查询中,查询的细微变化可能导致最佳匹配结果排名大幅波动。本文系统分析了多种非语义查询扰动在多媒体信息检索场景中的影响,使用TRECVID Ad-Hoc Video Search查询和V3C1视频数据集,评估了多种CLIP变体。结果显示,语法和语义扰动引起的不稳定性最大,而脆弱性主要集中在标点、大小写等表面修改上。研究强调,鲁棒性应成为评估视觉-语言模型的关键维度,超越基准准确率。
原文摘要 · Abstract (English)
Multimodal co-embedding models, especially CLIP, have advanced the state of the art in zero-shot classification and multimedia information retrieval in recent years by aligning images and text in a shared representation space. However, such modals trained on a contrastive alignment can lack stability towards small input perturbations. Especially when dealing with manually expressed queries, minor variations in the query can cause large differences in the ranking of the best-matching results. In this paper, we present a systematic analysis of the effect of multiple classes of non-semantic query perturbations in an multimedia information retrieval scenario. We evaluate a diverse set of lexical, syntactic, and semantic perturbations across multiple CLIP variants using the TRECVID Ad-Hoc Video Search queries and the V3C1 video collection. Across models, we find that syntactic and semantic perturbations drive the largest instabilities, while brittleness is concentrated in trivial surface edits such as punctuation and case. Our results highlight robustness as a critical dimension for evaluating vision-language models beyond benchmark accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。