用大模型自动检查科学图表是否符合可视化规范,发现常见错误。
Evaluating Compliance with Visualization Guidelines in Diagrams for Scientific Publications Using Large Vision Language Models
- 用大视觉语言模型分析图表,匹配多种可视化准则
- 能准确识别缺失标签、多余3D效果等问题,部分指标F1超96%
- 适合科研人员自查图表质量,提升论文可读性
图表广泛用于科学出版物中的数据可视化。数据可视化领域定义了创建和使用图表的原则与指南,但研究人员常不了解或未遵守,导致信息不准确或不完整。本文利用大视觉语言模型(VLMs)分析图表,识别其在选定数据可视化原则上的潜在问题。通过对比五种开源VLMs和五种提示策略,基于来自可视化指南的问题集进行评估。结果表明,所用VLM在识别图表类型(F1-score 82.49%)、3D效果(F1-score 98.55%)、坐标轴标签(F1-score 76.74%)、线条(RMSE 1.16)、颜色(RMSE 1.60)和图例(F1-score 96.64%,RMSE 0.70)方面表现良好,但在图像质量(F1-score 0.74%)和刻度标记/标签(F1-score 46.13%)方面不可靠。Qwen2.5VL在所有模型中表现最佳,总结式提示策略在多数问题上效果最优。研究证明VLM可自动识别大量图表问题,如缺失坐标轴标签、缺少图例和不必要的3D效果。该方法可扩展至数据可视化的其他方面。
原文摘要 · Abstract (English)
Diagrams are widely used to visualize data in publications. The research field of data visualization deals with defining principles and guidelines for the creation and use of these diagrams, which are often not known or adhered to by researchers, leading to misinformation caused by providing inaccurate or incomplete information. In this work, large Vision Language Models (VLMs) are used to analyze diagrams in order to identify potential problems in regards to selected data visualization principles and guidelines. To determine the suitability of VLMs for these tasks, five open source VLMs and five prompting strategies are compared using a set of questions derived from selected data visualization guidelines. The results show that the employed VLMs work well to accurately analyze diagram types (F1-score 82.49 %), 3D effects (F1-score 98.55 %), axes labels (F1-score 76.74 %), lines (RMSE 1.16), colors (RMSE 1.60) and legends (F1-score 96.64 %, RMSE 0.70), while they cannot reliably provide feedback about the image quality (F1-score 0.74 %) and tick marks/labels (F1-score 46.13 %). Among the employed VLMs, Qwen2.5VL performs best, and the summarizing prompting strategy performs best for most of the experimental questions. It is shown that VLMs can be used to automatically identify a number of potential issues in diagrams, such as missing axes labels, missing legends, and unnecessary 3D effects. The approach laid out in this work can be extended for further aspects of data visualization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。