用可验证奖励训练模型,让图表推理更准更可信。
Chart-RVR: Reinforcement Learning with Verifiable Rewards for Explainable Chart Reasoning
- 结合可验证奖励与强化学习,提升模型对图表的推理能力。
- 在分布内和分布外数据上均超越传统微调方法,准确率更高。
- 生成的推理过程更真实可信,适合需要解释性的应用场景。
大型视觉语言模型在图表推理等视觉推理任务中已达到顶尖水平,但在分布外(OOD)数据上表现仍差,且要求生成思维链(CoT)推理过程时性能进一步下降,限制了可解释性。我们提出 Chart-RVR 框架,通过将组相对策略优化(GRPO)与自动可验证奖励相结合,对 30 亿参数的 LVLM 进行微调,增强其在图表推理中的鲁棒性和可解释性。该框架包含三项奖励:(i) 正确的图表类型分类,(ii) 忠实的表格重建,(iii) 推理过程一致性。在分布内和分布外数据集上,Chart-RVR 均优于标准监督微调(SFT),缩小了 OOD 性能差距,并提升了推理内容的真实性。所获模型(Chart-RVR-3B 系列)在六个涵盖域内与分布外设置的图表推理基准上达到当前最优,超越同规模所有现有模型。除准确性外,生成的 CoT 推理更易理解,增强了可信度与可靠性,展现了可验证奖励与 GRPO 在训练可靠、可解释图表推理模型方面的强大潜力。
原文摘要 · Abstract (English)
The capabilities of Large Vision-Language Models (LVLMs) have reached state-of-the-art on many visual reasoning tasks, including chart reasoning, yet they still falter on out-of-distribution (OOD) data, and degrade further when asked to produce their chain-of-thought (CoT) rationales, limiting explainability. We present Chart-RVR, a general framework that fine-tunes LVLMs to be more robust and explainable for chart reasoning by coupling Group Relative Policy Optimization (GRPO) with automatically verifiable rewards. Our framework comprises of three rewards that maximize: (i) correct chart-type classification, (ii) faithful chart table reconstruction, and (iii) process conformity. Applied to 3-billion-parameter LVLMs, Chart-RVR consistently outperforms standard supervised fine-tuning (SFT) on both in-distribution and out-of-distribution datasets, closing the OOD performance gap while improving rationale fidelity. The resulting models, the Chart-RVR-3B series, achieve state-of-the-art results on six chart-reasoning benchmarks spanning in-domain and OOD settings, surpassing all existing models of comparable size. Beyond accuracy, Chart-RVR yields more interpretable CoT rationales, strengthening trust and reliability - showcasing the power of verifiable rewards with GRPO for training reliable, interpretable chart-reasoning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。