arXiv:2511.06805cs.AIcs.LG2025-11AAAI被引 2

通过自我演进的迭代反思,提升多模态模型解数学题能力。

MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning

  • 采用迭代推理+反思+奖励反馈循环优化模型
  • 在MathVL-test上超越领先开源模型QVQ
  • 适合需要强数学推理的多模态应用研究者

多模态大语言模型在视觉-语言问答任务中表现优异,但在复杂推理任务(如数学问题求解)中仍面临挑战。以往方法依赖从教师模型蒸馏的专用数学数据集,仅捕捉静态推理模式,难以适应新问题或更复杂题目,缺乏迭代深化能力。为此,我们提出 extbf{/method}——一种面向多模态模型的数学自我演进框架。不同于传统一次性微调,/method通过推理、反思与奖励引导反馈的多轮迭代机制,融合前序阶段的正确推理路径和专门设计的成果奖励模型(ORM)的反思反馈。在多个挑战性基准测试中验证了其有效性,实验结果表明,在MathVL-test上显著优于基线模型,性能超越当前领先的开源多模态数学推理模型QVQ。代码与模型已公开于https://zheny2751.com/MathSE.github.io/。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in vision-language answering tasks. Despite their strengths, these models often encounter challenges in achieving complex reasoning tasks such as mathematical problem-solving. Previous works have focused on fine-tuning on specialized mathematical datasets. However, these datasets are typically distilled directly from teacher models, which capture only static reasoning patterns and leaving substantial gaps compared to student models. This reliance on fixed teacher-derived datasets not only restricts the model's ability to adapt to novel or more intricate questions that extend beyond the confines of the training data, but also lacks the iterative depth needed for robust generalization. To overcome these limitations, we propose \textbf{\method}, a \textbf{Math}ematical \textbf{S}elf-\textbf{E}volving framework for MLLMs. In contrast to traditional one-shot fine-tuning paradigms, \method iteratively refines the model through cycles of inference, reflection, and reward-based feedback. Specifically, we leverage iterative fine-tuning by incorporating correct reasoning paths derived from previous-stage inference and integrating reflections from a specialized Outcome Reward Model (ORM). To verify the effectiveness of \method, we evaluate it on a suite of challenging benchmarks, demonstrating significant performance gains over backbone models. Notably, our experimental results on MathVL-test surpass the leading open-source multimodal mathematical reasoning model QVQ. Our code and models are available at \texttt{https://zheny2751\allowbreak-dotcom.github.io/\allowbreak MathSE.github.io/}.

数学推理多模态自我演进迭代优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。