构建细粒度评估基准CHIRP,提升视觉语言模型开放问答的评测可靠性。
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models
- 用多尺度组合LLM与视觉编码器构建Robin模型集,发现现有评测方法缺陷。
- 提出CHIRP长文本响应基准,覆盖更全面的评估场景。
- 开源Robin训练代码、模型和CHIRP数据集,推动研究可复现性。
近年来视觉语言模型(VLMs)快速发展,亟需严谨全面的评估方法与基准。本文分析了现有VLM评估技术,包括自动指标、AI辅助评估和人类评价,覆盖多种任务。首先,我们构建了Robin——一个通过多尺度融合大语言模型(LLMs)与视觉编码器(VEs)得到的新模型集合,并利用Robin揭示了当前评估方法在不同尺度上的不足。为克服这些局限,我们提出了新的长文本响应评估基准CHIRP,以实现更稳健、完整的VLM评估。我们已公开Robin的训练代码、模型套件及CHIRP基准数据,以促进可复现性并推进VLM研究。
原文摘要 · Abstract (English)
The proliferation of Vision-Language Models (VLMs) in the past several years calls for rigorous and comprehensive evaluation methods and benchmarks. This work analyzes existing VLM evaluation techniques, including automated metrics, AI-based assessments, and human evaluations across diverse tasks. We first introduce Robin - a novel suite of VLMs that we built by combining Large Language Models (LLMs) and Vision Encoders (VEs) at multiple scales, and use Robin to identify shortcomings of current evaluation approaches across scales. Next, to overcome the identified limitations, we introduce CHIRP - a new long form response benchmark we developed for more robust and complete VLM evaluation. We provide open access to the Robin training code, model suite, and CHIRP benchmark to promote reproducibility and advance VLM research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。