评测六款多模态大模型在科学可视化理解上的表现,发现模型能力参差不齐。
Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

- 构建49题科学可视化测试集,涵盖18个图示与8类可视化技术。
- Gemini表现最佳,部分任务超越人类平均,开源模型普遍低于人类基准。
- 模型在定量估算、流向判断等任务上频繁出错,适合评估多模态AI真实理解力。
多模态大语言模型(MLLMs)被广泛用于解释可视化内容,但现有评估仍以图表为中心,缺乏对科学可视化(SciVis)理解力的充分验证。本文在包含49个题目的标准化科学可视化素养测评中,对六款MLLM进行了基准测试,该测试基于18个科学可视化图示和插图,覆盖8种可视化技术与11类任务。评估采用闭域协议,对比了三款闭源与三款开源模型,结果基于485名人类参与者的数据。结果显示,当前MLLM在科学可视化理解上表现不一:Gemini整体最强,多个子集表现超过人类均值;而开源模型整体低于人类基准。模型在科学插图、搜索与空间理解任务中表现较好,但在纹理型与整合型可视化,以及定量估算任务中表现较差。错误分析揭示其在细粒度定量估计、流线方向解读与基于上下文的编码理解上存在系统性缺陷。研究建议将科学可视化素养作为评估多模态AI系统的必要维度。代码与模型输出已公开于https://github.com/patdmp/mllm-scivis-lit-benchmark。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We evaluate three closed-source and three open-source models under a closed-world protocol and compare their performance using data from 485 human participants. Results show that current MLLMs do not exhibit uniform SciVis literacy. Gemini is the strongest model overall, exceeding the human mean across the evaluated subsets, whereas the open-source models remain below the human baseline. Performance is highly uneven across techniques and tasks: models perform best on scientific illustration, search, and spatial understanding, but struggle on texture-based and integration-based visualizations and on quantitative estimation. Error analysis reveals recurring failures in fine-grained quantitative estimation, flow-direction interpretation, and grounded encoding interpretation. These findings position SciVis literacy as a necessary benchmark dimension for evaluating multimodal AI systems. Our code and model outputs are publicly available at https://github.com/patdmp/mllm-scivis-lit-benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。