arXiv:2603.27064cs.CVcs.AI2026-03中稿 · CVPR被引 4

构建百万级图表数据集,提升模型对图表的跨模态理解能力

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding

  • 用代码生成150万张多样化图表,包含24种类型和6个绘图库
  • 每张图表配图像、数据表、文本摘要和带推理的问答,实现细粒度对齐
  • 含真实数据与安全标注子集,适合训练鲁棒的可视化理解模型

理解图表需模型联合推理几何视觉模式、结构化数值数据与自然语言,而当前视觉语言模型仍有限。我们提出ChartNet,一个高质量、百万级规模的多模态数据集,用于推动图表解析与推理。ChartNet采用新型代码引导合成流程,生成150万种多样化的图表样本,涵盖24种图表类型与6个绘图库。每个样本包含五部分对齐内容:绘图代码、渲染图像、数据表格、自然语言摘要及带推理的问答,实现细粒度跨模态对齐。为覆盖完整图表理解场景,数据集还包含人工标注、真实世界数据、安全与定位子集。通过严格的质量过滤流程,确保图像保真度、语义准确性和图表表达多样性。在多个基准上微调均显著提升性能,验证其作为大规模监督信号的有效性。作为同类最大开源数据集,ChartNet旨在支持具备强鲁棒性与泛化能力的数据可视化理解基础模型发展。数据集已公开于https://huggingface.co/datasets/ibm-granite/ChartNet。

原文摘要 · Abstract (English)

Understanding charts requires models to jointly reason over geometric visual patterns, structured numerical data, and natural language -- a capability where current vision-language models (VLMs) remain limited. We introduce ChartNet, a high-quality, million-scale multimodal dataset designed to advance chart interpretation and reasoning. ChartNet leverages a novel code-guided synthesis pipeline to generate 1.5 million diverse chart samples spanning 24 chart types and 6 plotting libraries. Each sample consists of five aligned components: plotting code, rendered chart image, data table, natural language summary, and question-answering with reasoning, providing fine-grained cross-modal alignment. To capture the full spectrum of chart comprehension, ChartNet additionally includes specialized subsets encompassing human annotated data, real-world data, safety, and grounding. Moreover, a rigorous quality-filtering pipeline ensures visual fidelity, semantic accuracy, and diversity across chart representations. Fine-tuning on ChartNet consistently improves results across benchmarks, demonstrating its utility as large-scale supervision for multimodal models. As the largest open-source dataset of its kind, ChartNet aims to support the development of foundation models with robust and generalizable capabilities for data visualization understanding. The dataset is publicly available at https://huggingface.co/datasets/ibm-granite/ChartNet

多模态图表理解数据集视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。