让AI医生学会问诊与决策,更像真人医生。
Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement Learning
- 用多智能体环境+双层奖励机制训练AI,同时提升问诊策略和诊断准确率。
- 在多个评测集上超越开源与闭源模型,参数效率更高。
- 适合医疗AI研发、临床辅助系统设计者参考。
人类医生在门诊服务中的专业性依赖于两个核心能力:精准的医疗决策能力和富有策略性与同理心的患者问诊技巧。现有大语言模型在医疗决策基准上已取得显著准确率,但往往缺乏高效、共情的问诊能力,难以适应真实临床场景。为此,我们提出Doctor-R1,一个通过提出高价值问题并开展策略性多轮问诊来指导决策的AI医生代理。该框架包含三个关键组件:多智能体交互环境、分层优化诊疗与沟通技能的双层奖励架构,以及基于高质量历史轨迹的经验存储库。我们在OpenAI的HealthBench和MAQuE数据集上评估Doctor-R1,涵盖沟通质量、用户体验和任务准确率等多维度指标。结果显示,Doctor-R1以更高的参数效率显著超越当前最先进的开源专用大模型,并优于强大闭源模型。人类专家评估也表明,其在临床能力与以患者为中心的表现上均表现优异,验证了该框架的有效性。
原文摘要 · Abstract (English)
The professionalism of a human doctor in outpatient service depends on two core abilities: the ability to make accurate medical decisions and the medical consultation skill to conduct strategic, empathetic patient inquiry. Existing Large Language Models (LLMs) have achieved remarkable accuracy on medical decision-making benchmarks. However, they often lack the ability to conduct the strategic and empathetic consultation, which is essential for real-world clinical scenarios. To address this gap, we propose Doctor-R1, an AI doctor agent trained to master both of the capabilities by ask high-yield questions and conduct strategic multi-turn inquiry to guide decision-making. Our framework introduces three key components: a multi-agent interactive environment, a two-tiered reward architecture that separately optimizes clinical decision-making and communicative inquiry skills, and an experience repository to ground policy learning in high-quality prior trajectories. We evaluate Doctor-R1 on OpenAI's HealthBench and MAQuE, assessed across multi-facet metrics, such as communication quality, user experience, and task accuracy. Remarkably, Doctor-R1 surpasses state-of-the-art open-source specialized LLMs by a substantial margin with higher parameter efficiency and outperforms powerful proprietary models. Furthermore, the human expert evaluations show that Doctor-R1 achieves superior clinical capability and patient-centric performance, demonstrating the effectiveness of the framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。