arXiv:2506.12307cs.CLcs.AI2025-06被引 13

用强化学习统一提升大模型在各类医学问答中的推理能力

Med-U1: Incentivizing Unified Medical Reasoning in LLMs via Large-scale Reinforcement Learning

论文配图:Med-U1: Incentivizing Unified Medical Reasoning in LLMs via Large-scale Reinforcement Learning
图 1 · 摘自论文原文
  • 通过大规模强化学习与规则奖励,统一优化多种医学问答任务
  • 在多个医学基准上超越更大更专的模型,且输出更简洁可验证
  • 适合医疗AI研究者及需要高可靠性推理系统的开发者

医学问答(Med-QA)涵盖多项选择题、开放文本生成和复杂计算推理等多种任务。尽管近年来增强推理的大语言模型取得进展,但其全面医学理解能力仍不明确。本文提出Med-U1,一种统一框架,可在多格式输出任务(从选择题到复杂生成与计算)中实现稳健推理。Med-U1采用纯大规模强化学习,结合混合规则型二元奖励函数,并引入长度惩罚控制输出冗余。通过多目标奖励优化,引导模型生成简明且可验证的推理链。实证结果表明,Med-U1在多个挑战性医学问答基准上显著提升性能,甚至优于更大的专用及商用模型。此外,该方法对分布外(OOD)任务也表现出强泛化能力。深入分析揭示了训练策略、推理链长度控制与奖励设计的关键洞见。代码已公开。

原文摘要 · Abstract (English)

Medical Question-Answering (QA) encompasses a broad spectrum of tasks, including multiple choice questions (MCQ), open-ended text generation, and complex computational reasoning. Despite this variety, a unified framework for delivering high-quality medical QA has yet to emerge. Although recent progress in reasoning-augmented large language models (LLMs) has shown promise, their ability to achieve comprehensive medical understanding is still largely unexplored. In this paper, we present Med-U1, a unified framework for robust reasoning across medical QA tasks with diverse output formats, ranging from MCQs to complex generation and computation tasks. Med-U1 employs pure large-scale reinforcement learning with mixed rule-based binary reward functions, incorporating a length penalty to manage output verbosity. With multi-objective reward optimization, Med-U1 directs LLMs to produce concise and verifiable reasoning chains. Empirical results reveal that Med-U1 significantly improves performance across multiple challenging Med-QA benchmarks, surpassing even larger specialized and proprietary models. Furthermore, Med-U1 demonstrates robust generalization to out-of-distribution (OOD) tasks. Extensive analysis presents insights into training strategies, reasoning chain length control, and reward design for medical LLMs. Our code is available here.

医学推理强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。