评测8个视觉语言模型在6项人类数据可视化理解任务中的表现,发现模型普遍不如人类。
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
- 设计6项面向人类的可视化理解测试,评估模型认知能力。
- 模型平均表现低于人类,且在错误模式上与人类差异显著。
- 结果揭示当前模型在数据推理方面仍有巨大提升空间,适合研究人机认知差异者参考。
数据可视化是传达定量数据中模式的强大工具。然而,理解任何数据可视化都并非易事——这需要同时解析视觉、数值和语言输入,并按既定格式进行整合,而这种格式需通过先前学习才能掌握。近期发展的视觉语言模型理论上是构建此类认知行为计算模型的有力候选。但目前尚不清楚这些模型在涉及数据可视化推理的任务中,能多大程度模拟人类表现。这一差距源于以往研究使用与人类评估标准不同的度量方式。本文对8个视觉语言模型进行了评估,采用6项专为人类设计的数据可视化素养测试,并将模型输出与人类参与者对比。结果显示,这些模型平均表现劣于人类,即使采用较宽松的标准,差距依然存在。此外,尽管模型与人类在各题目上的相对表现存在一定相关性,但所有模型的错误模式均显著区别于人类。综合来看,当前人工系统在模拟人类数据可视化推理方面仍有巨大改进空间。所有代码和数据均可在 https://osf.io/e25mu/?view_only=399daff5a14d4b16b09473cf19043f18 获取。
原文摘要 · Abstract (English)
Data visualizations are powerful tools for communicating patterns in quantitative data. Yet understanding any data visualization is no small feat -- succeeding requires jointly making sense of visual, numerical, and linguistic inputs arranged in a conventionalized format one has previously learned to parse. Recently developed vision-language models are, in principle, promising candidates for developing computational models of these cognitive operations. However, it is currently unclear to what degree these models emulate human behavior on tasks that involve reasoning about data visualizations. This gap reflects limitations in prior work that has evaluated data visualization understanding in artificial systems using measures that differ from those typically used to assess these abilities in humans. Here we evaluated eight vision-language models on six data visualization literacy assessments designed for humans and compared model responses to those of human participants. We found that these models performed worse than human participants on average, and this performance gap persisted even when using relatively lenient criteria to assess model performance. Moreover, while relative performance across items was somewhat correlated between models and humans, all models produced patterns of errors that were reliably distinct from those produced by human participants. Taken together, these findings suggest significant opportunities for further development of artificial systems that might serve as useful models of how humans reason about data visualizations. All code and data needed to reproduce these results are available at: https://osf.io/e25mu/?view_only=399daff5a14d4b16b09473cf19043f18.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。