arXiv:2505.08468cs.CLcs.CV2025-05ACL被引 8

用开源大模型当评委,自动评估图表理解能力。

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?

  • 设计成对与单点评估任务,测试13个开源模型做评委的能力。
  • 部分模型达GPT-4水平(80%一致率),但有模型仅10%一致率。
  • 适合研究与工业界低成本自动化评测图表任务的场景。

图表广泛用于辅助人们理解与推理数据。近期涌现出图表问答、图表转文本、事实核查等下游任务。大视觉语言模型(LVLMs)在这些任务上展现潜力,但其评估成本高、耗时长,限制了实际部署。若使用LVLM作为评委来评估其他LVLM的图表理解能力,可简化流程,但受限于专有数据集、模型访问权限及评估成本,难以在工业界推广。为此,我们全面评估了13个开源LVLM在多种图表理解与推理任务中的评委表现。设计了成对与单点评估任务,涵盖事实正确性、信息量和相关性等标准。还分析了格式遵循、位置一致性、长度偏好与指令遵循能力。聚焦参数量小于10B的低成本模型,采用标准化协议与评分体系,衡量评委准确性。实验结果揭示显著差异:部分开源模型达成约80%与GPT-4判断的一致性,达到GPT-4级表现;另一些模型则低于10%一致率。研究显示,当前最先进的开源LVLM可作为图表任务的低成本自动评估工具,但位置偏好与长度偏差等问题仍存在。

原文摘要 · Abstract (English)

Charts are ubiquitous as they help people understand and reason with data. Recently, various downstream tasks, such as chart question answering, chart2text, and fact-checking, have emerged. Large Vision-Language Models (LVLMs) show promise in tackling these tasks, but their evaluation is costly and time-consuming, limiting real-world deployment. While using LVLMs as judges to assess the chart comprehension capabilities of other LVLMs could streamline evaluation processes, challenges like proprietary datasets, restricted access to powerful models, and evaluation costs hinder their adoption in industrial settings. To this end, we present a comprehensive evaluation of 13 open-source LVLMs as judges for diverse chart comprehension and reasoning tasks. We design both pairwise and pointwise evaluation tasks covering criteria like factual correctness, informativeness, and relevancy. Additionally, we analyze LVLM judges based on format adherence, positional consistency, length bias, and instruction-following. We focus on cost-effective LVLMs (<10B parameters) suitable for both research and commercial use, following a standardized evaluation protocol and rubric to measure the LVLM judge's accuracy. Experimental results reveal notable variability: while some open LVLM judges achieve GPT-4-level evaluation performance (about 80% agreement with GPT-4 judgments), others struggle (below ~10% agreement). Our findings highlight that state-of-the-art open-source LVLMs can serve as cost-effective automatic evaluators for chart-related tasks, though biases such as positional preference and length bias persist.

图表理解大模型评测开源模型自动评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。