arXiv:2608.08623cs.AI2026-08被引 1

医学数学推理中用知识引导的奖励框架,提升准确性与安全性。

MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning

论文配图:MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning
图 1 · 摘自论文原文
  • 引入知识验证机制,强制生成可解释的计算公式。
  • 混合软硬奖励策略,兼顾临床安全阈值与精度敏感性。
  • 在医疗场景下显著提升推理准确率与泛化能力,适合安全关键领域。

在可验证奖励的强化学习(RLVR)框架中,数学推理任务通常采用基于容差的浮点结果评估方式,但该方法在临床场景下面临阈值校准困难、训练不稳定和精度有限等问题。为此,我们提出一种知识引导的混合奖励框架(MedCalc-R1)。具体而言,引入知识验证奖励机制,强制模型显式生成计算公式,并通过外部验证器进行检验,以增强可解释性与推理可靠性。此外,设计了一种结合硬约束(基于临床安全阈值)与软奖励(精度敏感)的混合策略,逐步引导模型在安全范围内优化。实验表明,该方法在推理准确率和泛化能力上显著优于现有基线,在安全关键领域具有有效性和适用性。

原文摘要 · Abstract (English)

In Reinforcement Learning with Verifiable Rewards (RLVR) frameworks for mathematical reasoning tasks, floating-point results are typically evaluated using a tolerance-based reward. However, this strategy suffers from challenges such as difficulty in threshold calibration, unstable training dynamics, and limited accuracy, especially in clinical scenarios. To address these limitations, we propose a knowledge-guided hybrid reward framework (\textsc{MedCalc-R1}). Specifically, we introduce a knowledge verification reward mechanism that enforces explicit generation of computational formulas, which are further validated by an external verifier to enhance interpretability and reasoning reliability. Furthermore, we design a hybrid soft-hard reward scheme combining a hard constraint based on clinical safety thresholds with a soft, precision-sensitive reward that progressively guides learning within the acceptable range. Experimental results demonstrate that our method significantly outperforms existing baselines in both reasoning accuracy and generalization capability, validating the effectiveness and applicability in safety-critical domains.

医学推理强化学习知识引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。