arXiv:2505.15265cs.CVcs.AI2025-05

用进化算法找让视觉语言模型出错的敏感语义,提升模型鲁棒性。

Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs

  • 结合LLM与文生图模型,通过演化搜索敏感语义概念。
  • 在7个主流LVLM上验证,成功发现导致错误的关键语义。
  • 适合研究模型安全与可解释性的学者参考。

对抗攻击旨在生成误导深度模型的恶意输入,但无法提供可解释信息,如‘输入中哪些内容导致模型易出错’。这一信息对提升模型鲁棒性至关重要。近期研究发现,模型对某些视觉语义(如‘潮湿’、‘雾天’)特别敏感,易产生错误。本文首次探索大视觉语言模型(LVLMs)的敏感语义,发现其在特定语义概念下易出现幻觉和错误。为此,我们提出一种新颖的语义演化框架:利用大语言模型(LLM)和文生图(T2I)模型,让随机初始化的语义概念通过LLM驱动的交叉与变异生成图像描述,并由T2I模型生成视觉输入。LVLM在任务上的表现作为语义的适应度分数,反馈给LLM以引导更优语义的探索。在7个主流LVLM和2个多模态任务上的实验验证了方法有效性。此外,我们揭示了LVLM对敏感语义的规律性响应,为后续研究提供启发。

原文摘要 · Abstract (English)

Adversarial attacks aim to generate malicious inputs that mislead deep models, but beyond causing model failure, they cannot provide certain interpretable information such as ``\textit{What content in inputs make models more likely to fail?}'' However, this information is crucial for researchers to specifically improve model robustness. Recent research suggests that models may be particularly sensitive to certain semantics in visual inputs (such as ``wet,'' ``foggy''), making them prone to errors. Inspired by this, in this paper we conducted the first exploration on large vision-language models (LVLMs) and found that LVLMs indeed are susceptible to hallucinations and various errors when facing specific semantic concepts in images. To efficiently search for these sensitive concepts, we integrated large language models (LLMs) and text-to-image (T2I) models to propose a novel semantic evolution framework. Randomly initialized semantic concepts undergo LLM-based crossover and mutation operations to form image descriptions, which are then converted by T2I models into visual inputs for LVLMs. The task-specific performance of LVLMs on each input is quantified as fitness scores for the involved semantics and serves as reward signals to further guide LLMs in exploring concepts that induce LVLMs. Extensive experiments on seven mainstream LVLMs and two multimodal tasks demonstrate the effectiveness of our method. Additionally, we provide interesting findings about the sensitive semantics of LVLMs, aiming to inspire further in-depth research.

模型安全语义敏感演化搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。