构建百万级信息图图表数据集,提升模型理解与生成复杂图表能力
ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation
- 基于真实图表归纳75种类型、440种变体、68种版式,程序化生成合成数据
- 在微调后显著提升模型对信息图图表的理解性能,支持代码生成与示例驱动生成
- 适合研究多模态推理、图表生成与设计自动化的人群使用
信息图图表通过结合视觉元素(如图表、图像)与文本信息,成为传达抽象数据的有力媒介。然而其丰富的视觉与结构特征给大视觉语言模型(LVLMs)带来挑战,因为这些模型通常仅在简单图表上训练。为填补这一差距,我们提出ChartGalaxy,一个百万规模的数据集,旨在推动信息图图表的理解与生成。该数据集通过归纳法从真实信息图中识别出75种图表类型、440种图表变体和68种版式布局,并程序化生成合成数据。我们展示了该数据集的实用性:1)通过微调提升信息图图表理解能力;2)建立信息图图表代码生成基准;3)实现基于示例的信息图图表生成。ChartGalaxy通过捕捉真实设计中的视觉与结构复杂性,为增强LVLMs的多模态推理与生成能力提供了宝贵资源。
原文摘要 · Abstract (English)
Infographic charts are a powerful medium for communicating abstract data by combining visual elements (e.g., charts, images) with textual information. However, their visual and structural richness poses challenges for large vision-language models (LVLMs), which are typically trained on plain charts. To bridge this gap, we introduce ChartGalaxy, a million-scale dataset designed to advance the understanding and generation of infographic charts. The dataset is constructed through an inductive process that identifies 75 chart types, 440 chart variations, and 68 layout templates from real infographic charts and uses them to create synthetic ones programmatically. We showcase the utility of this dataset through: 1) improving infographic chart understanding via fine-tuning, 2) benchmarking code generation for infographic charts, and 3) enabling example-based infographic chart generation. By capturing the visual and structural complexity of real design, ChartGalaxy provides a useful resource for enhancing multimodal reasoning and generation in LVLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。