构建新基准测试大模型看图推理能力,发现其远逊于人类。
ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
- 设计仅靠视觉推理可解的合成数据集,暴露模型短板。
- 真实世界图表问答集含1162题,人类准确率93%,最佳模型仅63%。
- 视觉推理任务性能下降35%-55%,适合评估视觉理解能力的模型。
图表理解对大型视觉语言模型(LVLM)构成独特挑战,需融合复杂文本与视觉推理能力。但现有模型在视觉推理方面明显弱于文本推理。我们通过一个仅依赖视觉推理即可解答的合成数据集进行案例研究,发现随着视觉复杂度增加,模型性能显著下降,而人类表现保持稳健。随后提出ChartMuseum,一个包含1,162个专家标注问题的新图表问答基准,覆盖多种推理类型,数据源自184个真实来源。与以往基准不同,该基准下模型与人类表现差距显著:人类准确率达93%,最佳模型Gemini-2.5-Pro仅达63.0%,领先开源模型Qwen2.5-VL-72B-Instruct为38.5%。在纯视觉推理任务中,所有模型性能相比文本主导任务下降35%-55%。定性分析揭示当前模型在特定视觉推理类别上存在系统性困难。
原文摘要 · Abstract (English)
Chart understanding presents a unique challenge for large vision-language models (LVLMs), as it requires the integration of sophisticated textual and visual reasoning capabilities. However, current LVLMs exhibit a notable imbalance between these skills, falling short on visual reasoning that is difficult to perform in text. We conduct a case study using a synthetic dataset solvable only through visual reasoning and show that model performance degrades significantly with increasing visual complexity, while human performance remains robust. We then introduce ChartMuseum, a new Chart Question Answering (QA) benchmark containing 1,162 expert-annotated questions spanning multiple reasoning types, curated from real-world charts across 184 sources, specifically built to evaluate complex visual and textual reasoning. Unlike prior chart understanding benchmarks -- where frontier models perform similarly and near saturation -- our benchmark exposes a substantial gap between model and human performance, while effectively differentiating model capabilities: although humans achieve 93% accuracy, the best-performing model Gemini-2.5-Pro attains only 63.0%, and the leading open-source LVLM Qwen2.5-VL-72B-Instruct achieves only 38.5%. Moreover, on questions requiring primarily visual reasoning, all models experience a 35%-55% performance drop from text-reasoning-heavy question performance. Lastly, our qualitative error analysis reveals specific categories of visual reasoning that are challenging for current LVLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。