构建合成数据集评估大模型看图能力,涵盖五类图表。
PUB: Plot Understanding Benchmark and Dataset for Evaluating Large Language Models on Synthetic Visual Data Interpretation
- 用可控参数生成合成图表数据,确保覆盖真实场景。
- 测试多模型对图表的解读准确率,发现差异性表现。
- 适合研究视觉理解、人机交互与自动化分析的学者。
大型语言模型(LLMs)理解数据可视化表示的能力对于推动其在数据分析和决策中的应用至关重要。本文提出一个新型合成数据集,用于评估LLMs在解读时间序列、直方图、小提琴图、箱线图和聚类图等各类图表方面的表现。数据通过受控参数生成,全面覆盖现实世界可能遇到的场景。采用包含视觉数据的多模态文本提示进行评测,考察ChatGPT、Gemini等先进模型的理解与解析准确性。为保证数据完整性,本基准数据集全自动生成,完全未被待测模型接触过,从而避免预训练响应干扰,实现无偏评估。我们引入定量指标,提供一套稳健且全面的评估工具。对多个前沿模型的基准测试显示其在不同图表类型上的表现存在显著差异,揭示了当前模型的优势与短板。研究结果为理解现有模型能力提供了关键洞见,并指明改进方向。该工作为未来提升语言模型视觉解读能力的研究奠定了基础。具备强视觉理解能力的改进型模型可广泛应用于自动数据分析、科研辅助、教育工具及商业智能等领域。
原文摘要 · Abstract (English)
The ability of large language models (LLMs) to interpret visual representations of data is crucial for advancing their application in data analysis and decision-making processes. This paper presents a novel synthetic dataset designed to evaluate the proficiency of LLMs in interpreting various forms of data visualizations, including plots like time series, histograms, violins, boxplots, and clusters. Our dataset is generated using controlled parameters to ensure comprehensive coverage of potential real-world scenarios. We employ multimodal text prompts with questions related to visual data in images to benchmark several state-of-the-art models like ChatGPT or Gemini, assessing their understanding and interpretative accuracy. To ensure data integrity, our benchmark dataset is generated automatically, making it entirely new and free from prior exposure to the models being tested. This strategy allows us to evaluate the models' ability to truly interpret and understand the data, eliminating possibility of pre-learned responses, and allowing for an unbiased evaluation of the models' capabilities. We also introduce quantitative metrics to assess the performance of the models, providing a robust and comprehensive evaluation tool. Benchmarking several state-of-the-art LLMs with this dataset reveals varying degrees of success, highlighting specific strengths and weaknesses in interpreting diverse types of visual data. The results provide valuable insights into the current capabilities of LLMs and identify key areas for improvement. This work establishes a foundational benchmark for future research and development aimed at enhancing the visual interpretative abilities of language models. In the future, improved LLMs with robust visual interpretation skills can significantly aid in automated data analysis, scientific research, educational tools, and business intelligence applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。