用自训练生成合成图表数据,提升模型真实场景下的图表理解能力。
EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart Understanding
- 自训练生成高质量合成图表数据,同时优化模型性能
- 新基准EvoChart-QA含650张真实网页图表与1250个专家问题
- 开源模型在真实图表上准确率提升至54.2%,逼近闭源模型水平
图表理解可实现人类的自动化数据分析,要求模型具备高精度的视觉理解能力。尽管现有视觉语言模型(VLMs)在图表理解方面取得进展,但高质量训练数据匮乏和全面评估基准的缺失仍阻碍其发展。本文提出EvoChart,一种用于生成合成图表数据的自训练方法,以增强VLM在真实场景中的图表理解能力;同时构建EvoChart-QA,一个全新的评估基准。该基准包含从140个不同网站收集的650张真实世界图表,以及1250个由专家精心设计的问题。在多个开源与专有VLM上测试的结果表明,即使是最先进的专有模型GPT-4o,在EvoChart-QA上的准确率也仅为49.8%。而EvoChart方法显著提升了开源VLM在真实图表理解任务中的表现,达到54.2%的准确率。
原文摘要 · Abstract (English)
Chart understanding enables automated data analysis for humans, which requires models to achieve highly accurate visual comprehension. While existing Visual Language Models (VLMs) have shown progress in chart understanding, the lack of high-quality training data and comprehensive evaluation benchmarks hinders VLM chart comprehension. In this paper, we introduce EvoChart, a novel self-training method for generating synthetic chart data to enhance VLMs' capabilities in real-world chart comprehension. We also propose EvoChart-QA, a noval benchmark for measuring models' chart comprehension abilities in real-world scenarios. Specifically, EvoChart is a unique self-training data synthesis approach that simultaneously produces high-quality training corpus and a high-performance chart understanding model. EvoChart-QA consists of 650 distinct real-world charts collected from 140 different websites and 1,250 expert-curated questions that focus on chart understanding. Experimental results on various open-source and proprietary VLMs tested on EvoChart-QA demonstrate that even the best proprietary model, GPT-4o, achieves only 49.8% accuracy. Moreover, the EvoChart method significantly boosts the performance of open-source VLMs on real-world chart understanding tasks, achieving 54.2% accuracy on EvoChart-QA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。