arXiv:2501.12051cs.CL2025-01AAAI被引 20

用自进化框架让小模型具备精准医疗推理能力

MedS$^3$: Towards Medical Slow Thinking with Self-Evolved Soft Dual-sided Process Supervision

  • 通过自探索生成可验证的推理路径,逐步优化小模型
  • 在11个基准上比现有最佳医疗模型高6.45分,超32B大模型8.57分
  • 适合需要可靠、轻量级医疗推理的应用场景

医学语言模型在真实临床推理应用中面临重大挑战。当前主流方法在任务覆盖、中间推理步骤的细粒度监督以及依赖专有系统方面存在不足,难以实现通用、可信且高效的临床推理模型。为此,我们提出MedS3,一个自进化框架,赋予小型可部署模型强大的推理能力。基于从五个医学领域和16个数据集中采样的8000条精选实例,利用小规模基础策略模型进行蒙特卡洛树搜索(MCTS),构建规则可验证的推理轨迹。根据节点值排序的自探索推理路径用于通过强化微调和偏好学习来增强策略模型。此外,引入软双过程奖励模型,结合价值动态:降低节点值的步骤将被惩罚,从而即使最终答案正确也能精确定位推理错误。在11个基准上的实验表明,MedS3相比之前最先进的医疗模型提升6.45个百分点,超越320亿参数的一般推理模型8.57个百分点。额外实证分析进一步证明,MedS3实现了稳健且忠实的推理行为。

原文摘要 · Abstract (English)

Medical language models face critical barriers to real-world clinical reasoning applications. However, mainstream efforts, which fall short in task coverage, lack fine-grained supervision for intermediate reasoning steps, and rely on proprietary systems, are still far from a versatile, credible and efficient language model for clinical reasoning usage. To this end, we propose MedS3, a self-evolving framework that imparts robust reasoning capabilities to small, deployable models. Starting with 8,000 curated instances sampled via a curriculum strategy across five medical domains and 16 datasets, we use a small base policy model to conduct Monte Carlo Tree Search (MCTS) for constructing rule-verifiable reasoning trajectories. Self-explored reasoning trajectories ranked by node values are used to bootstrap the policy model via reinforcement fine-tuning and preference learning. Moreover, we introduce a soft dual process reward model that incorporates value dynamics: steps that degrade node value are penalized, enabling fine-grained identification of reasoning errors even when the final answer is correct. Experiments on eleven benchmarks show that MedS3 outperforms the previous state-of-the-art medical model by +6.45 accuracy points and surpasses 32B-scale general-purpose reasoning models by +8.57 points. Additional empirical analysis further demonstrates that MedS3 achieves robust and faithful reasoning behavior.

医疗推理自进化小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。