用真实图表生成数据集+强化学习,提升模型看图分析能力。
BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
- 通过真实图表重绘生成多样且准确的视觉数据集。
- 结合监督微调与强化学习,在多个评测中超越现有模型。
- 适合需要精准图表理解的科研、金融等场景使用。
图表在数据分析中至关重要,将原始数据转化为直观视觉表达以辅助决策。尽管当前视觉-语言模型(VLMs)取得显著进展,仍因训练数据缺乏多样性与真实感,或依赖自动提取的含误差数据表而难以理解图表。现有方法仅使用低质量数据进行监督微调,严重限制效果。为此,我们提出BigCharts数据集构建流程,通过真实来源图表条件化渲染生成视觉多样化的图表图像。相比纯合成数据集,BigCharts保留真实数据并确保视觉多样性,同时通过重构绘制过程保证底层数据准确性。此外,我们设计了融合监督微调与基于组相对策略优化(GRPO)的强化学习的综合训练框架,引入专为图表推理设计的新奖励信号,显著提升模型在不同图表风格与领域下的鲁棒性与泛化能力,得到当前最优的图表推理模型BigCharts-R1。大量实验表明,该模型在多个图表问答基准上超越现有方法,包括更大规模的开源与闭源模型。
原文摘要 · Abstract (English)
Charts are essential to data analysis, transforming raw data into clear visual representations that support human decision-making. Although current vision-language models (VLMs) have made significant progress, they continue to struggle with chart comprehension due to training on datasets that lack diversity and real-world authenticity, or on automatically extracted underlying data tables of charts, which can contain numerous estimation errors. Furthermore, existing models only rely on supervised fine-tuning using these low-quality datasets, severely limiting their effectiveness. To address these issues, we first propose BigCharts, a dataset creation pipeline that generates visually diverse chart images by conditioning the rendering process on real-world charts sourced from multiple online platforms. Unlike purely synthetic datasets, BigCharts incorporates real-world data, ensuring authenticity and visual diversity, while still retaining accurate underlying data due to our proposed replotting process. Additionally, we introduce a comprehensive training framework that integrates supervised fine-tuning with Group Relative Policy Optimization (GRPO)-based reinforcement learning. By introducing novel reward signals specifically designed for chart reasoning, our approach enhances model robustness and generalization across diverse chart styles and domains, resulting in a state-of-the-art chart reasoning model, BigCharts-R1. Extensive experiments demonstrate that our models surpass existing methods on multiple chart question-answering benchmarks compared to even larger open-source and closed-source models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。