arXiv:2607.02995cs.CVcs.AI2026-07

检测视觉语言模型在特定图像概念下的异常响应模式。

VISTA: Auditing Semantic Divergence in Vision-Language Models

论文配图:VISTA: Auditing Semantic Divergence in Vision-Language Models
图 1 · 摘自论文原文
  • 用语义熵与分布差异结合,跨模型黑盒审计
  • 发现142个高可疑案例,部分模型对特定人群拒绝率高达65%
  • 揭示了模型对身份相关查询的选择性拒绝新问题

视觉语言模型在面对包含人口统计特征、企业标志或意识形态符号的图像时,可能表现出视觉概念依赖的响应偏差:某些模型生成异常一致的回答,与同类模型存在显著差异。这类行为难以通过纯文本审计发现,因为视觉概念无法像文本词元那样被隔离或替换。我们提出VISTA(视觉不一致性筛查分析),一种基于语义熵与分布差异的黑盒跨模型审计方法,可识别模型特异性异常。在受控实验中,通过对三个VLM在小规模偏见数据集上微调植入概念相关立场,验证了VISTA的有效性。在涵盖19个主题的六款VLM审计中,VISTA共发现142个高可疑案例(占1.2%),并首次识别出‘选择性拒绝’这一新型偏差模式——不同群体的询问被拒绝率从0%至65%不等。

原文摘要 · Abstract (English)

Vision-language models can exhibit visual concept-conditioned divergence: given images containing demographic features, corporate logos, or ideological symbols, some models produce unusually uniform responses that differ from what peer models say about the same input. These behaviors evade text-only audits because visual concepts cannot be isolated or substituted the way text tokens can. We present VISTA (Visual Inconsistency Screening Through Analysis), a black-box cross-model audit that couples semantic entropy with distribution-based divergence to flag model-specific anomalies. In a controlled study, we implant concept-conditioned stances in three VLMs via fine-tuning on small biased datasets and confirm that VISTA detects them. Auditing six VLMs across 19 topics, VISTA surfaces 142 high-suspicion cases (1.2%) and identifies selective refusal as a previously unreported divergence pattern, where models refuse demographic queries at rates varying from 0 to 65% across groups.

视觉语言模型偏差检测黑盒审计语义熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。