arXiv:2512.00601cs.AI2025-12被引 4

用多目标强化学习让大模型看病更准确、更可信、更全面

Clinical-R1: Empowering Large Language Models for Faithful and Comprehensive Reasoning with Clinical Objective Relative Policy Optimization

  • 设计可验证的奖励机制,同时优化准确率、忠实度和完整性
  • 在3个医学基准上,比传统方法提升真实性与完整性,准确率也略有提高
  • 无需人工标注,适合医疗等高风险领域的大模型训练

大语言模型(LLM)通过大规模预训练和后训练强化学习展现出强大推理能力,如DeepSeek-R1。然而,现有后训练方法(如GRPO)主要奖励正确性,与医学等高风险领域所需的多维目标不一致,临床推理还需保证忠实性和完整性。本文提出临床目标相对策略优化(CRPO),一种可扩展的多目标、可验证强化学习方法,通过规则驱动和可验证奖励信号联合优化准确性、忠实性和完整性,无需人工标注。我们训练了30亿参数的Clinical-R1-3B模型用于临床推理。在三个基准测试中,相比标准GRPO,CRPO显著提升推理的真实性与完整性,同时保持适度的准确率增益。该框架为对齐大模型推理与临床目标提供了可扩展路径,推动更安全、协作的医疗AI发展,并揭示了多目标、可验证强化学习在医学领域大模型后训练中的潜力。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have shown strong reasoning capabilities through large-scale pretraining and post-training reinforcement learning, demonstrated by DeepSeek-R1. However, current post-training methods, such as Grouped Relative Policy Optimization (GRPO), mainly reward correctness, which is not aligned with the multi-dimensional objectives required in high-stakes fields such as medicine, where reasoning must also be faithful and comprehensive. We introduce Clinical-Objective Relative Policy Optimization (CRPO), a scalable, multi-objective, verifiable reinforcement learning method designed to align LLM post-training with clinical reasoning principles. CRPO integrates rule-based and verifiable reward signals that jointly optimize accuracy, faithfulness, and comprehensiveness without relying on human annotation. To demonstrate its effectiveness, we train Clinical-R1-3B, a 3B-parameter model for clinical reasoning. The experiments on three benchmarks demonstrate that our CRPO substantially improves reasoning on truthfulness and completeness over standard GRPO while maintaining comfortable accuracy enhancements. This framework provides a scalable pathway to align LLM reasoning with clinical objectives, enabling safer and more collaborative AI systems for healthcare while also highlighting the potential of multi-objective, verifiable RL methods in post-training scaling of LLMs for medical domains.

大模型推理医疗AI强化学习可信生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。