arXiv:2503.18432cs.CLcs.AI2025-03被引 2

用强化学习让大模型学会逐步批改数学题,提升推理能力。

Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning

  • 将逐步批改转化为强化学习问题,增强模型推理能力。
  • 在两个数据集上超越11个基线模型,准确率显著提升。
  • 适合教育AI、自动评分系统研究者参考。

自动数学批改旨在通过人工智能技术检查学生解答数学问题的过程。现有研究多关注问题级别的最终答案判断,而忽视了对解题每一步的详细反馈,这需要语义理解与推理能力。本文提出一种基于强化学习的改进方法StepAMC,将文本分类任务中的逐步自动批改转化为强化学习问题,以增强大语言模型的推理能力。设计了空间受限的策略网络以提高强化学习的稳定性,并引入细粒度奖励网络,将二值人类反馈转换为连续奖励值。在两个基准数据集上进行大量实验,结果表明,所提模型优于11个强基线模型。

原文摘要 · Abstract (English)

Automatic math correction aims to check students' solutions to mathematical problems via artificial intelligence technologies. Most existing studies focus on judging the final answer at the problem level, while they ignore detailed feedback on each step in a math problem-solving process, which requires abilities of semantic understanding and reasoning. In this paper, we propose a reinforcement learning (RL)-based method to boost large language model (LLM) for step-level automatic math correction, named StepAMC. Particularly, we convert the step-level automatic math correction within the text classification task into an RL problem to enhance the reasoning capabilities of LLMs. Then, we design a space-constrained policy network to improve the stability of RL. Then, we introduce a fine-grained reward network to convert the binary human feedback into a continuous value. We conduct extensive experiments over two benchmark datasets and the results show that our model outperforms the eleven strong baselines.

数学教育强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。