构建可扩展的视觉语言模型评估框架,实现低成本高精度领域专用测试集生成。
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
- 通过任务增强技术,从一个任务生成多个多样化子任务
- 发布7个领域共16万+条人工验证答案的统一标准基准集
- 在3.7万任务上评测22个SOTA模型,揭示跨领域性能差异
可靠评估对AI模型的科学发展和实际应用至关重要。现有视觉语言模型(VLM)基准存在设计异质、覆盖领域有限等问题,难以支持跨域比较与特定领域评估。为此,本文提出三项贡献:(1) 基于任务增强的资源高效生成领域专用VLM基准框架;(2) 发布7个新领域基准,遵循统一协议,包含162,946条经过严格人工验证的答案;(3) 在总计37,171个任务上对22个前沿VLM进行大规模评测,揭示模型在不同领域与任务间的显著性能差异,证明定制化基准的必要性。该方法有助于资源高效地选择适配特定场景的模型,并推动未来研究聚焦核心开放问题。
原文摘要 · Abstract (English)
Reliable evaluation of AI models is critical for scientific progress and practical application. While existing VLM benchmarks provide general insights into model capabilities, their heterogeneous designs and limited focus on a few imaging domains pose significant challenges for both cross-domain performance comparison and targeted domain-specific evaluation. To address this, we propose three key contributions: (1) a framework for the resource-efficient creation of domain-specific VLM benchmarks enabled by task augmentation for creating multiple diverse tasks from a single existing task, (2) the release of new VLM benchmarks for seven domains, created according to the same homogeneous protocol and including 162,946 thoroughly human-validated answers, and (3) an extensive benchmarking of 22 state-of-the-art VLMs on a total of 37,171 tasks, revealing performance variances across domains and tasks, thereby supporting the need for tailored VLM benchmarks. Adoption of our methodology will pave the way for the resource-efficient domain-specific selection of models and guide future research efforts toward addressing core open questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。