用零样本提示评估视觉语言模型对图表的感知能力
Understanding Graphical Perception in Data Visualization through Zero-shot Prompting of Vision-Language Models
- 用零样本提示测试VLM在图表理解任务中的表现
- 部分任务和风格下,VLM表现接近人类水平
- 图表颜色填充和连续性会影响VLM判断,类似人类
视觉语言模型(VLM)在需要结合图表图像与文本描述的图表理解任务中表现优异。然而,其性能是否具备类人认知行为尚不明确。若能证明VLM具备类人图表理解能力,则可拓展用于可视化设计与评估。本文通过零样本提示,评估VLM在具有成熟人类表现基准的图形感知任务上的准确性。结果表明,在特定任务与风格组合下,VLM表现与人类相当,具备建模人类表现的潜力。此外,输入样式的变化显示,即使数据与映射关系不变,填色方式和图表连续性也会显著影响VLM准确率,说明其对视觉风格敏感,与人类感知机制一致。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) have been successful at many chart comprehension tasks that require attending to both the images of charts and their accompanying textual descriptions. However, it is not well established how VLM performance profiles map to human-like behaviors. If VLMs can be shown to have human-like chart comprehension abilities, they can then be applied to a broader range of tasks, such as designing and evaluating visualizations for human readers. This paper lays the foundations for such applications by evaluating the accuracy of zero-shot prompting of VLMs on graphical perception tasks with established human performance profiles. Our findings reveal that VLMs perform similarly to humans under specific task and style combinations, suggesting that they have the potential to be used for modeling human performance. Additionally, variations to the input stimuli show that VLM accuracy is sensitive to stylistic changes such as fill color and chart contiguity, even when the underlying data and data mappings are the same.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。