研究生成式视觉语言模型对提示词语义和词汇变化的敏感性
Sensitivity of Generative VLMs to Semantically and Lexically Altered Prompts
- 用SugarCrepe++数据集测试提示词微调后模型对词汇/语义变化的响应
- 发现生成式VLM对无语义变化的词汇改动极度敏感
- 揭示提示一致性技术易受干扰,适合关注模型鲁棒性的研究者
尽管生成式视觉语言模型(VLMs)的提示调优技术大量涌现,但其对提示词中词汇与语义变化的敏感性仍不明确。本文利用SugarCrepe++数据集,评估生成式VLMs理解提示词中词汇与语义变化的能力。分析了在无对应语义改变的情况下,提示词词汇变化对模型的影响。结果表明,生成式VLMs对此类变化高度敏感。此外,该脆弱性影响了旨在实现输出一致性的技术性能。
原文摘要 · Abstract (English)
Despite the significant influx of prompt-tuning techniques for generative vision-language models (VLMs), it remains unclear how sensitive these models are to lexical and semantic alterations in prompts. In this paper, we evaluate the ability of generative VLMs to understand lexical and semantic changes in text using the SugarCrepe++ dataset. We analyze the sensitivity of VLMs to lexical alterations in prompts without corresponding semantic changes. Our findings demonstrate that generative VLMs are highly sensitive to such alterations. Additionally, we show that this vulnerability affects the performance of techniques aimed at achieving consistency in their outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。