arXiv:2508.04650cs.CV2025-08被引 11

新基准EncQA揭示视觉模型在图表理解上的认知短板。

EncQA: Benchmarking Vision-Language Models on Visual Encodings for Charts

  • 构建覆盖6类编码与8类任务的合成问答数据集
  • 9个主流模型表现差异大,大模型不必然更优
  • 适合研究图表推理与视觉认知的学者参考

多模态视觉语言模型在图表理解任务上持续取得进展,但现有成果未能全面反映图表解读所需的视觉推理能力。本文提出EncQA,一个基于可视化研究文献的新基准,系统覆盖六种视觉编码(位置、长度、面积、定量色彩、名义色彩、形状)和八类分析任务(找极值、查数值、找异常、筛选值、计算精确派生值、计算相对派生值、相关性判断、相对相关性判断),包含2076组合成问答对。对9个先进VLM的评估显示,同一任务中不同编码的表现差异显著,且多数任务-编码组合下模型规模增大并未带来性能提升。结果表明,推动图表理解需针对性弥补特定视觉推理缺口,而非单纯扩大模型或数据规模。

原文摘要 · Abstract (English)

Multimodal vision-language models (VLMs) continue to achieve ever-improving scores on chart understanding benchmarks. Yet, we find that this progress does not fully capture the breadth of visual reasoning capabilities essential for interpreting charts. We introduce EncQA, a novel benchmark informed by the visualization literature, designed to provide systematic coverage of visual encodings and analytic tasks that are crucial for chart understanding. EncQA provides 2,076 synthetic question-answer pairs, enabling balanced coverage of six visual encoding channels (position, length, area, color quantitative, color nominal, and shape) and eight tasks (find extrema, retrieve value, find anomaly, filter values, compute derived value exact, compute derived value relative, correlate values, and correlate values relative). Our evaluation of 9 state-of-the-art VLMs reveals that performance varies significantly across encodings within the same task, as well as across tasks. Contrary to expectations, we observe that performance does not improve with model size for many task-encoding pairs. Our results suggest that advancing chart understanding requires targeted strategies addressing specific visual reasoning gaps, rather than solely scaling up model or dataset size.

图表理解视觉推理多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。