用可验证奖励强化学习,让AI更懂图表中的数学逻辑。
Chart-RL: Generalized Chart Comprehension via Reinforcement Learning with Verifiable Rewards
- 设计可数学验证的奖励机制,引导模型精准推理图表数据。
- 在多图表问答任务上提升16.7%,复杂任务仅需10个样本即超6000个简单样本训练。
- 擅长处理视觉变化,适合需要强推理能力的图表理解场景。
准确理解图表是推动多模态学习系统发展的关键挑战,因大量信息被压缩为结构化视觉表示。现有视觉语言模型(VLMs)在未见图表上泛化能力差,因其需对结构化视觉内容进行抽象、符号与定量推理。本文提出Chart-RL,一种基于可数学验证奖励的强化学习方法,用于提升VLM在图表问答任务中的表现。实验表明,Chart-RL在多个图表理解基准上持续优于监督微调(SFT),在MutlChartQA上相对提升16.7%,在ChartInsights上提升11.5%。鲁棒性分析显示,其在25种扰动图表类别中,有18类性能更优,体现强一致性和推理能力。此外,任务难度与内在复杂性比数据量更重要:仅用10个复杂图表-查询样例训练的Chart-RL,显著超越使用超过6000个简单样本训练的模型。且在困难推理任务上训练,不仅提升域内泛化,还促进向域外视觉数学问题的强迁移。
原文摘要 · Abstract (English)
Accurate chart comprehension represents a critical challenge in advancing multimodal learning systems, as extensive information is compressed into structured visual representations. However, existing vision-language models (VLMs) frequently struggle to generalize on unseen charts because it requires abstract, symbolic, and quantitative reasoning over structured visual representations. In this work, we introduce Chart-RL, an effective reinforcement learning (RL) method that employs mathematically verifiable rewards to enhance chart question answering in VLMs. Our experiments demonstrate that Chart-RL consistently outperforms supervised fine-tuning (SFT) across different chart understanding benchmarks, achieving relative improvements of 16.7% on MutlChartQA, and 11.5% on ChartInsights. We conduct robustness analysis, where Chart-RL achieves enhanced performance in 18 of 25 perturbed chart categories, demonstrating strong consistency and reasoning capability across visual variations. Furthermore, we demonstrate that task difficulty and inherent complexity are more critical than data quantity in RL training. For instance, Chart-RL trained on merely 10 complex chart-query examples significantly outperforms models trained on over 6,000 simple examples. Additionally, training on challenging reasoning tasks not only improves in-domain generalization relative to simpler tasks, but also facilitate strong transfer to out-of-domain visual mathematical problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。