arXiv:2602.18094cs.CVcs.AI2026-02被引 1

为大模型设计的分布外数据评估基准,揭示其在真实场景下的脆弱性。

OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models

  • 构建4万条分布外图像-类别对,自动标注减少人工干预。
  • 现有大模型在常见类别上仍出现显著性能下降。
  • 提出渐进式提问评估法,更全面衡量复杂度影响。

现有视觉语言模型(VLMs)在大规模数据集上取得显著进展,通常假设数据独立同分布(IID)。但在真实场景中,这一假设常不成立。若无法妥善处理分布外(OOD)对象,可能引发安全风险(如自动驾驶或医疗辅助)。然而,当前研究缺乏有效基准来全面评估VLMs在处理OOD数据时的表现。为此,我们提出OODBench,一种主要自动化、人工验证极少的新基准构建与评估方法。该基准包含4万条实例级的分布外实例-类别对。实验表明,即使底层图像类别常见,当前VLMs在OODBench上仍表现出显著性能下降。此外,我们提出一种可靠的自动化评估指标,通过基础到高级的提示问题序列,更充分评估不同难度问题受OOD数据的影响。最后,我们总结了大量发现与洞见,以促进未来对分布外数据获取与评估的研究。

原文摘要 · Abstract (English)

Existing Visual-Language Models (VLMs) have achieved significant progress by being trained on massive-scale datasets, typically under the assumption that data are independent and identically distributed (IID). However, in real-world scenarios, it is often impractical to expect that all data processed by an AI system satisfy this assumption. Furthermore, failure to appropriately handle out-of-distribution (OOD) objects may introduce safety risks in real-world applications (e.g., autonomous driving or medical assistance). Unfortunately, current research has not yet provided valid benchmarks that can comprehensively assess the performance of VLMs in response to OOD data. Therefore, we propose OODBench, a predominantly automated method with minimal human verification, for constructing new benchmarks and evaluating the ability of VLMs to process OOD data. OODBench contains 40K instance-level OOD instance-category pairs, and we show that current VLMs still exhibit notable performance degradation on OODBench, even when the underlying image categories are common. In addition, we propose a reliable automated assessment metric that employs a Basic-to-Advanced Progression of prompted questions to assess the impact of OOD data on questions of varying difficulty more fully. Lastly, we summarize substantial findings and insights to facilitate future research in the acquisition and evaluation of OOD data.

视觉语言模型分布外检测评估基准AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。