arXiv:2511.19032cs.CV2025-11被引 2

提出新评测框架Bench-C,揭示视觉噪声下VLM模型的隐藏可靠性问题。

Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models

  • 构建可控测试集Bench-C,覆盖19种噪声类型与5级严重度
  • 发现轻微噪声可提升准确率但破坏预测结构,存在隐性退化
  • 引入RAS指标追踪信心与正确性对齐,适合关注模型鲁棒性的研究者

视觉噪声可能以顶1准确率无法捕捉的方式改变视觉-语言模型(VLM)的行为。模型可能保持相同答案,但分布支持已丧失;或通过不稳定的错误到正确转变提升准确率。我们提出Bench-C,一个受控的多项选择测试平台,选取语义多样、对噪声敏感的样本,在19种噪声类型和5个严重度级别下评估模型表现。为衡量噪声如何改变选项分布,我们引入鲁棒性对齐分数(RAS),结合置信度-正确性对齐与不确定性方向。进一步区分原始正确与原始错误样本,追踪变化是否为临时或持续。在13个VLM上的实验揭示反直觉现象:轻微噪声可提升顶1准确率,却损害预测结构。此类失效包括沉默退化、错误自信及严重度依赖的持续性。Bench-C因此支持超越最终答案的鲁棒性评估,并定位可靠性变化发生位置。代码与数据见https://github.com/xiangjieSui/Bench-C。

原文摘要 · Abstract (English)

Visual corruptions can change vision--language model (VLM) behavior in ways that top-1 accuracy does not capture. A model may keep the same answer while losing distributional support, or improve accuracy through unstable wrong-to-correct changes. We introduce Bench-C, a controlled multiple-choice testbed for studying these effects. It selects semantically diverse samples whose predictions respond to corruption, and evaluates them under 19 corruption types and five severity levels. To measure how corruption changes the option distribution, we introduce the Robustness Alignment Score (RAS), which combines confidence-correctness alignment with uncertainty direction. We further separate originally correct samples from originally wrong samples, and track whether changes are temporary or persistent across severity. Experiments across 13 VLMs reveal a counterintuitive pattern: mild corruptions can improve top-1 accuracy while degrading prediction structure. These failures include silent degradation, erroneous overconfidence, and severity-dependent persistence. Bench-C therefore supports robustness evaluation that goes beyond final answers and attributes where reliability changes occur. Code and data are available at https://github.com/xiangjieSui/Bench-C.

模型鲁棒性视觉语言模型可靠性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。