全面评估视觉语言模型的9个关键维度,揭示效率模型在偏见问题上的短板。
VHELM: A Holistic Evaluation of Vision Language Models
- 构建覆盖9大维度的统一评估框架,涵盖感知、推理、公平性等
- 22个模型在21个数据集上测试,发现轻量模型在偏见任务中表现显著更差
- 提供自动化、可复现的评测流程,适合研究者对比模型综合性能
当前视觉语言模型(VLMs)的评测基准多聚焦于感知或解题能力,忽视了公平性、多语言性、毒性等关键方面。且评估方式和范围各异,难以横向比较。为此,我们扩展HELM框架,提出视觉语言模型的全面评估体系(VHELM),整合多个数据集,覆盖视觉感知、知识、推理、偏见、公平性、多语言性、鲁棒性、毒性与安全共9个维度。通过标准化推理参数、提示方法和评估指标,实现模型间的公平比较。框架轻量自动,评测高效。首次评估涵盖22个VLMs在21个现有数据集上的表现,发现以效率为导向的模型(如Claude 3 Haiku、Gemini 1.5 Flash)在偏见基准上显著劣于全尺寸模型(如Claude 3 Opus、Gemini 1.5 Pro),但在其他维度差异较小。所有原始生成结果与完整数据已公开于官网(https://crfm.stanford.edu/helm/vhelm/v2.0.1)。VHELM将持续更新,欢迎持续贡献。
原文摘要 · Abstract (English)
Current benchmarks for assessing vision-language models (VLMs) often focus on their perception or problem-solving capabilities and neglect other critical aspects such as fairness, multilinguality, or toxicity. Furthermore, they differ in their evaluation procedures and the scope of the evaluation, making it difficult to compare models. To address these issues, we extend the HELM framework to VLMs to present the Holistic Evaluation of Vision Language Models (VHELM). VHELM aggregates various datasets to cover one or more of the 9 aspects: visual perception, knowledge, reasoning, bias, fairness, multilinguality, robustness, toxicity, and safety. In doing so, we produce a comprehensive, multi-dimensional view of the capabilities of the VLMs across these important factors. In addition, we standardize the standard inference parameters, methods of prompting, and evaluation metrics to enable fair comparisons across models. Our framework is designed to be lightweight and automatic so that evaluation runs are cheap and fast. Our initial run evaluates 22 VLMs on 21 existing datasets to provide a holistic snapshot of the models. We uncover new key findings, such as the fact that efficiency-focused models (e.g., Claude 3 Haiku or Gemini 1.5 Flash) perform significantly worse than their full models (e.g., Claude 3 Opus or Gemini 1.5 Pro) on the bias benchmark but not when evaluated on the other aspects. For transparency, we release the raw model generations and complete results on our website (https://crfm.stanford.edu/helm/vhelm/v2.0.1). VHELM is intended to be a living benchmark, and we hope to continue adding new datasets and models over time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。