用偏好排序优化提升医学问答的推理准确性和临床可靠性。
RPRO: Ranked Preference Reinforcement Optimization for Enhancing Medical QA and Diagnostic Reasoning
- 基于偏好排序与强化学习,自动识别并修正低质量推理链。
- 20亿参数模型在多项数据集上超越70亿至200亿参数大模型。
- 适合医疗AI研发、临床辅助系统构建者参考使用。
医学问答需要整合领域知识与逻辑推理的高级能力。然而,现有大语言模型生成的推理链常缺乏事实准确性与临床可靠性。本文提出排名偏好强化优化(RPRO)框架,结合强化学习与偏好驱动的推理优化,提升临床思维链(CoT)性能。RPRO采用任务自适应推理模板和概率评估机制,使模型输出更符合临床工作流程,并能自动识别与修正低质量推理链。不同于传统成对偏好方法,RPRO引入基于Bradley-Terry模型的组间排序优化,并加入KL散度正则化以保证训练稳定。在PubMedQA、MedQA-USMLE及台湾远东纪念医院(FEMH)真实临床数据集上的实验表明,RPRO持续优于强基线模型。值得注意的是,其20亿参数模型在多项指标上超越70亿至200亿参数的大型模型,包括专用于医学的变体。结果表明,将偏好优化与质量驱动的精炼相结合,是一种可扩展且具备临床基础的构建可靠医学大模型的新路径。
原文摘要 · Abstract (English)
Medical question answering requires advanced reasoning that integrates domain knowledge with logical inference. However, existing large language models (LLMs) often generate reasoning chains that lack factual accuracy and clinical reliability. We propose Ranked Preference Reinforcement Optimization (RPRO), a novel framework that combines reinforcement learning with preference-driven reasoning refinement to enhance clinical chain-of-thought (CoT) performance. RPRO distinguishes itself from prior approaches by employing task-adaptive reasoning templates and a probabilistic evaluation mechanism that aligns model outputs with established clinical workflows, while automatically identifying and correcting low-quality reasoning chains. Unlike traditional pairwise preference methods, RPRO introduces a groupwise ranking optimization based on the Bradley--Terry model and incorporates KL-divergence regularization for stable training. Experiments on PubMedQA, MedQA-USMLE, and a real-world clinical dataset from Far Eastern Memorial Hospital (FEMH) demonstrate consistent improvements over strong baselines. Remarkably, our 2B-parameter model outperforms much larger 7B--20B models, including medical-specialized variants. These findings demonstrate that combining preference optimization with quality-driven refinement provides a scalable and clinically grounded approach to building more reliable medical LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。