让AI从多种解题思路中学习,提升数学推理能力。
Multimodal Mathematical Reasoning with Diverse Solving Perspective
- 构建多解路径数据集,捕捉不同思考方式。
- 模型在数学基准上准确率与解法多样性双提升。
- 适合研究多视角推理与智能教育系统的开发者。
大规模强化学习的进展显著提升了大语言模型在数学领域中的推理能力。然而,现有的多模态大语言模型在数学推理中通常依赖一对一的图文配对和单一解法监督,忽略了有效解法的多样性与内部反思。本文提出MathV-DP数据集,为每个图像-问题对标注多种不同的解题路径,提供更丰富的推理监督。基于Qwen-VL模型,我们构建Qwen-VL-DP,通过监督学习微调并采用基于规则的强化学习方法——组相对策略优化(GRPO),融合正确性判断与多样性感知奖励函数。该方法强调从多样化推理视角学习,并区分正确但不同的解法。在MathVista minitest与Math-V基准上的实验表明,Qwen-VL-DP在准确率和生成多样性方面均显著优于现有基础多模态模型,凸显了引入多样解题视角与反思性推理的重要性。
原文摘要 · Abstract (English)
Recent progress in large-scale reinforcement learning (RL) has notably enhanced the reasoning capabilities of large language models (LLMs), especially in mathematical domains. However, current multimodal LLMs (MLLMs) for mathematical reasoning often rely on one-to-one image-text pairs and single-solution supervision, overlooking the diversity of valid reasoning perspectives and internal reflections. In this work, we introduce MathV-DP, a novel dataset that captures multiple diverse solution trajectories for each image-question pair, fostering richer reasoning supervision. We further propose Qwen-VL-DP, a model built upon Qwen-VL, fine-tuned with supervised learning and enhanced via group relative policy optimization (GRPO), a rule-based RL approach that integrates correctness discrimination and diversity-aware reward functions. Our method emphasizes learning from varied reasoning perspectives and distinguishing between correct yet distinct solutions. Extensive experiments on the MathVista's minitest and Math-V benchmarks demonstrate that Qwen-VL-DP significantly outperforms prior base MLLMs in both accuracy and generative diversity, highlighting the importance of incorporating diverse perspectives and reflective reasoning in multimodal mathematical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。