用可验证医学题训练大模型,提升复杂医疗推理能力。
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
- 构建可验证医学问题与验证器,指导模型生成正确推理路径。
- 仅用4万道题即超越通用和医学专用基线模型。
- 强化学习结合验证奖励,显著提升复杂医疗推理效果。
OpenAI o1 的突破表明增强推理能力可提升大模型表现,但现有研究多集中于数学任务,医疗等专业领域仍被忽视。医疗推理需高可靠性,但验证难度远高于数学。为此,我们提出可验证的医学问题及医学验证器,用于判断模型输出正确性。该可验证性支持两阶段方法:(1) 利用验证器引导搜索复杂推理轨迹以微调大模型;(2) 采用基于验证器奖励的强化学习进一步优化复杂推理。最终,我们推出 HuatuoGPT-o1,一款具备复杂医疗推理能力的大模型,仅使用40,000个可验证问题即超越通用和医学专用基线模型。实验表明,复杂推理显著提升医疗问题解决能力,且更受益于强化学习。本方法有望推动医疗及其他专业领域推理能力的发展。
原文摘要 · Abstract (English)
The breakthrough of OpenAI o1 highlights the potential of enhancing reasoning to improve LLM. Yet, most research in reasoning has focused on mathematical tasks, leaving domains like medicine underexplored. The medical domain, though distinct from mathematics, also demands robust reasoning to provide reliable answers, given the high standards of healthcare. However, verifying medical reasoning is challenging, unlike those in mathematics. To address this, we propose verifiable medical problems with a medical verifier to check the correctness of model outputs. This verifiable nature enables advancements in medical reasoning through a two-stage approach: (1) using the verifier to guide the search for a complex reasoning trajectory for fine-tuning LLMs, (2) applying reinforcement learning (RL) with verifier-based rewards to enhance complex reasoning further. Finally, we introduce HuatuoGPT-o1, a medical LLM capable of complex reasoning, which outperforms general and medical-specific baselines using only 40K verifiable problems. Experiments show complex reasoning improves medical problem-solving and benefits more from RL. We hope our approach inspires advancements in reasoning across medical and other specialized domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。