arXiv:2507.15509cs.AIcs.CV2025-07被引 20

用强化学习提升图表理解能力,让AI更准地分析复杂图表数据。

Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner

  • 通过分步推理生成高质量训练数据,覆盖多种图表类型。
  • 在多个基准上表现优于现有方法,接近大型闭源模型水平。
  • 适合需要精准图表分析的科研、金融与数据分析场景。

图表推理因其内在复杂性而面临独特挑战——需精确的数值理解、多层级视觉感知以及跨关联数据元素的逻辑推断。现有视觉语言模型在处理多子图场景和数值敏感任务时表现不佳。为此,我们提出 Chart-R1,一种面向图表领域的视觉语言模型,采用强化微调实现高级图表推理。首先,设计程序化数据合成方法,生成具有可验证答案格式的高质量分步推理数据,涵盖多样图表类型与复杂度。采用两阶段训练策略:(1) Chart-COT,通过链式思维监督将复杂推理分解为可解释的子任务;(2) Chart-RFT,使用群体相对策略优化并结合数值敏感奖励,针对图表特定推理进行优化。在开源基准和自建的 ChartRQA 数据集上的实验表明,Chart-R1 显著优于现有图表领域方法,且媲美大型开源/闭源模型。

原文摘要 · Abstract (English)

Chart reasoning presents unique challenges due to its inherent complexity -- requiring precise numerical comprehension, multi-level visual understanding, and logical inference across interconnected data elements. Existing vision-language models often struggle with such reasoning tasks, particularly when handling multi-subchart scenarios and numerical sensitivity. To address these challenges, we introduce Chart-R1, a chart-domain vision-language model that leverages reinforcement fine-tuning for advanced chart reasoning. We first propose a programmatic data synthesis approach to generate high-quality step-by-step reasoning data with verifiable answer formats, covering diverse chart types and complexity levels. Our two-stage training strategy includes: (1) Chart-COT, which decomposes complex reasoning into interpretable subtasks through chain-of-thought supervision, and (2) Chart-RFT, which employs group relative policy optimization with numerically sensitive rewards tailored for chart-specific reasoning. Experiments on open-source benchmarks and our proposed ChartRQA dataset demonstrate that Chart-R1 significantly outperforms existing chart-domain methods and rivals large-scale open/closed-source models.

图表推理强化学习链式思维视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。