arXiv:2409.18764cs.CVcs.CL2024-09被引 5

用视觉问答模型自动评估大模型生成图表的准确性与表达力。

Charting the Future: Using Chart Question-Answering for Scalable Evaluation of LLM-Driven Data Visualizations

  • 用VQA模型评测大模型生成图表的数据准确性和表达清晰度。
  • 大模型生成图表在VQA任务中表现低于人工绘制图表,但少样本提示显著提升精度。
  • 适合关注大模型可视化能力、评估方法创新的研究者使用。

我们提出一种新框架,利用视觉问答(VQA)模型自动化评估大语言模型生成的数据可视化效果。传统评估依赖人力,成本高且难以扩展;或仅关注数据准确性,忽略视觉传达效果。通过VQA模型,我们同时评估图表的数据呈现质量和整体沟通清晰度。实验基于ChartQA和PlotQA两个主流VQA基准数据集,使用OpenAI的GPT-3.5 Turbo和Meta的Llama 3.1 70B-Instruct生成图表。结果显示,大模型生成图表在VQA性能指标上未达人工生成图表水平;尽管少样本提示能显著提升生成准确性,但距离完全匹配人工绘图仍存明显差距。本工作通过无需人工标注的快速迭代机制,加速了该领域研究进展。

原文摘要 · Abstract (English)

We propose a novel framework that leverages Visual Question Answering (VQA) models to automate the evaluation of LLM-generated data visualizations. Traditional evaluation methods often rely on human judgment, which is costly and unscalable, or focus solely on data accuracy, neglecting the effectiveness of visual communication. By employing VQA models, we assess data representation quality and the general communicative clarity of charts. Experiments were conducted using two leading VQA benchmark datasets, ChartQA and PlotQA, with visualizations generated by OpenAI's GPT-3.5 Turbo and Meta's Llama 3.1 70B-Instruct models. Our results indicate that LLM-generated charts do not match the accuracy of the original non-LLM-generated charts based on VQA performance measures. Moreover, while our results demonstrate that few-shot prompting significantly boosts the accuracy of chart generation, considerable progress remains to be made before LLMs can fully match the precision of human-generated graphs. This underscores the importance of our work, which expedites the research process by enabling rapid iteration without the need for human annotation, thus accelerating advancements in this field.

图表生成VQA评估大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。