用专家标注的推理过程奖励模型替代投票,让医疗大模型更懂临床正确路径
MAPLE: Elevating Medical Reasoning from Statistical Consensus to Process-Led Alignment
- 用医学过程奖励模型替代传统多数投票,实现精准推理引导
- 在4个基准上均显著超越现有TTRL和单独使用PRM的方法
- 适合医疗AI研发者、需要可解释性推理的临床决策系统
近期医疗大模型研究探索了测试时强化学习(TTRL)以提升推理能力。然而,标准TTRL通常依赖多数投票(MV)作为启发式监督信号,在复杂医疗场景中,最频繁的推理路径未必是临床正确的。本文提出一种新统一训练范式,将医学过程奖励模型(Med-RPM)与TTRL结合,弥合测试时扩展(TTS)与参数化模型优化之间的差距。具体而言,我们通过细粒度、专家对齐的监督机制替代传统多数投票,使强化学习以医学正确性为导向,而非仅依赖共识。该方法有效将基于搜索的智能提炼进模型参数记忆中。在四个不同基准上的广泛评估表明,所提方法始终显著优于现有TTRL及独立使用PRM的选择策略。研究结果表明,从随机启发式转向结构化、步骤式奖励,是构建可靠且可扩展医疗AI系统的关键。
原文摘要 · Abstract (English)
Recent advances in medical large language models have explored Test-Time Reinforcement Learning (TTRL) to enhance reasoning. However, standard TTRL often relies on majority voting (MV) as a heuristic supervision signal, which can be unreliable in complex medical scenarios where the most frequent reasoning path is not necessarily the clinically correct one. In this work, we propose a novel and unified training paradigm that integrates medical process reward models with TTRL to bridge the gap between test-time scaling (TTS) and parametric model optimization. Specifically, we advance the TTRL framework by replacing the conventional MV with a fine-grained, expert-aligned supervision paradigm using Med-RPM. This integration ensures that reinforcement learning is guided by medical correctness rather than mere consensus, effectively distilling search-based intelligence into the model's parametric memory. Extensive evaluations on four different benchmarks have demonstrated that our developed method consistently and significantly outperforms current TTRL and standalone PRM selection. Our findings establish that transitioning from stochastic heuristics to structured, step-wise rewards is essential for developing reliable and scalable medical AI systems
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。