用大模型自动生成海量图表数据,训练出超越GPT-4V的图表理解模型。
SynChart: Synthesizing Charts from Language Models
- 仅用大语言模型生成多样化图表与密集标注数据
- 4.2B模型在ChartQA上接近GPT-4O表现,超越GPT-4V
- 适合需要高质量图表理解能力的研究者与开发者
随着GPT-4V(O)的发布,其在多模态任务中生成伪标签的应用日益流行。然而,如何从基础大语言模型构建此类先进模型仍不明确。本文探索仅使用大语言模型进行数据生成的潜力,并开发出专注于图表理解的高性能多模态模型。我们构建了大规模图表数据集SynChart,包含约400万张多样化的图表图像,以及超过7500万条密集标注,涵盖数据表、代码、描述和问答对。基于该数据集,我们训练了一个4.2亿参数的图表专家模型,在ChartQA任务上达到接近GPT-4O的性能,超越GPT-4V。结果表明,仅依赖语言模型即可实现高质量多模态模型的构建。
原文摘要 · Abstract (English)
With the release of GPT-4V(O), its use in generating pseudo labels for multi-modality tasks has gained significant popularity. However, it is still a secret how to build such advanced models from its base large language models (LLMs). This work explores the potential of using LLMs alone for data generation and develop competitive multi-modality models focusing on chart understanding. We construct a large-scale chart dataset, SynChart, which contains approximately 4 million diverse chart images with over 75 million dense annotations, including data tables, code, descriptions, and question-answer sets. We trained a 4.2B chart-expert model using this dataset and achieve near-GPT-4O performance on the ChartQA task, surpassing GPT-4V.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。