提出更真实、精准、高效的视觉语言模型评测方法
DatBench: Discriminative, Faithful, and Efficient VLM Evaluations
- 将选择题转为生成任务,避免猜测干扰
- 过滤掉70%的无关题目和42%错误样本,提升区分度
- 新基准提速13倍,适合大规模模型持续评估
实证评估是指导基础模型研究进展的主要指南。尽管大量工作聚焦于训练前沿视觉语言模型(VLMs),但其评估方法仍不成熟。为此,我们提出三个理想评估标准:(1) 与模态和应用场景忠实一致,(2) 能有效区分不同质量的模型,(3) 计算高效。通过此视角,我们识别出关键失败模式:(i) 选择题鼓励猜测,不反映下游应用,且模型性能提升后迅速饱和;(ii) 无需图像即可回答的问题占某些评估的高达70%;(iii) 错标或模糊样本在某些数据集中占比达42%。在效率方面,评估前沿模型的计算负担已极为沉重:有数据显示近20%的研发算力用于评估本身。我们并未废弃现有基准,而是通过转换与筛选进行优化,以最大化评估的保真度与区分度。将选择题改为生成任务后,模型性能下降最高达35%。同时,过滤掉无关和错标样本,在提升区分能力的同时降低计算成本。我们发布DatBench-Full——涵盖九种VLM能力的33个数据集清理版,以及更高效的DatBench子集,平均提速13倍(最高50倍),且与原始数据集的区分能力高度一致。本工作为未来模型持续扩展下的评估实践提供了严谨且可持续的路径。
原文摘要 · Abstract (English)
Empirical evaluation serves as the primary compass guiding research progress in foundation models. Despite a large body of work focused on training frontier vision-language models (VLMs), approaches to their evaluation remain nascent. To guide their maturation, we propose three desiderata that evaluations should satisfy: (1) faithfulness to the modality and application, (2) discriminability between models of varying quality, and (3) efficiency in compute. Through this lens, we identify critical failure modes that violate faithfulness and discriminability, misrepresenting model capabilities: (i) multiple-choice formats reward guessing, poorly reflect downstream use cases, and saturate early as models improve; (ii) blindly solvable questions, which can be answered without images, constitute up to 70% of some evaluations; and (iii) mislabeled or ambiguous samples compromise up to 42% of examples in certain datasets. Regarding efficiency, the computational burden of evaluating frontier models has become prohibitive: by some accounts, nearly 20% of development compute is devoted to evaluation alone. Rather than discarding existing benchmarks, we curate them via transformation and filtering to maximize fidelity and discriminability. We find that converting multiple-choice questions to generative tasks reveals sharp capability drops of up to 35%. In addition, filtering blindly solvable and mislabeled samples improves discriminative power while simultaneously reducing computational cost. We release DatBench-Full, a cleaned evaluation suite of 33 datasets spanning nine VLM capabilities, and DatBench, a discriminative subset that achieves 13x average speedup (up to 50x) while closely matching the discriminative power of the original datasets. Our work outlines a path toward evaluation practices that are both rigorous and sustainable as VLMs continue to scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。