arXiv:2505.15046cs.CLcs.AI2025-05被引 4

用统一框架生成图表元数据,让一张图支持多种任务。

ChartCards: A Chart-Metadata Generation Framework for Multi-Task Chart Understanding

  • 构建图表元数据框架,整合表格、代码、视觉元素等信息。
  • 生成17万条高质量图文描述,支持6种图表任务。
  • 在文本查图和图转表任务上提升超20%,适合多任务研究者。

多模态大模型为图表理解带来新机遇,但细粒度任务通常需大量高质量数据进行特定微调,成本高昂。为此,我们提出ChartCards——一种统一的图表元数据生成框架,系统合成包括数据表、可视化代码、视觉元素和多维语义描述在内的多种图表信息。通过结构化组织这些元数据,单张图表可支持文本到图表检索、图表摘要、图表转表格、图表描述及图表问答等多种下游任务。基于此,我们构建了大规模高质量数据集MetaChart,包含10,862个数据表、85,000张图表和170,000条高质量图表描述。通过众包评估与定量微调实验验证,将六种不同模型在MetaChart上微调后,各项任务平均性能提升5%;其中文本到图表检索和图表转表格任务提升最为显著,分别达到17%和28%(分别对应Long-CLIP和Llama 3.2-11B)。

原文摘要 · Abstract (English)

The emergence of Multi-modal Large Language Models (MLLMs) presents new opportunities for chart understanding. However, due to the fine-grained nature of these tasks, applying MLLMs typically requires large, high-quality datasets for task-specific fine-tuning, leading to high data collection and training costs. To address this, we propose ChartCards, a unified chart-metadata generation framework for multi-task chart understanding. ChartCards systematically synthesizes various chart information, including data tables, visualization code, visual elements, and multi-dimensional semantic captions. By structuring this information into organized metadata, ChartCards enables a single chart to support multiple downstream tasks, such as text-to-chart retrieval, chart summarization, chart-to-table conversion, chart description, and chart question answering. Using ChartCards, we further construct MetaChart, a large-scale high-quality dataset containing 10,862 data tables, 85K charts, and 170 K high-quality chart captions. We validate the dataset through qualitative crowdsourcing evaluations and quantitative fine-tuning experiments across various chart understanding tasks. Fine-tuning six different models on MetaChart resulted in an average performance improvement of 5% across all tasks. The most notable improvements are seen in text-to-chart retrieval and chart-to-table tasks, with Long-CLIP and Llama 3.2-11B achieving improvements of 17% and 28%, respectively.

图表理解多模态数据集构建元数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。