用专家反馈训练AI,让普通大模型也能精准诊断罕见病。
Teaching agentic AI to learn expert reasoning for rare disease diagnosis
- 通过人类反馈迭代策略,将专家推理转化为可复用的诊断政策。
- 在1243个病例中正确疾病排第一的比例从35.4%提升至59.3%。
- 政策可跨模型迁移且保持医生可控,适合医疗AI研发与临床辅助。
罕见病诊断依赖稀缺且难以传递的专家推理;现成大语言模型(LLMs)仅在35.4%的基准案例中将正确疾病排在首位。本文展示,通过受控学习流程而非单纯模型训练,可将专家推理转化为可扩展的AI能力。我们开发了liteOdyssey,一种基于人类反馈的策略迭代(PIHF)方法,该方法源自强化学习中的广义策略迭代,在此过程中模型失败与专家修正共同构建出由临床医生监管的策略,使现成LLM转变为自主诊断系统。结果表明,该策略显著提升诊断准确率,达到顶尖系统水平,部署开销极小,能泛化至未见疾病、跨模型迁移,并保持医生控制。在涵盖722种罕见病的1,243个公开测试案例中,正确疾病首列率从26.5%升至59.3%,在1,193个排除于策略训练外的案例与679种疾病上仍取得近似提升。消融实验显示,收益超过自动提示优化和源模型访问单独作用,且策略无需修改即可跨封闭与开源模型转移。在515名未确诊疾病网络患者中,liteOdyssey同样提升准确性,盲评医师更常认为其鉴别诊断精确,较少视为无用。通过PIHF,专家推理成为可被专家审查、修订并跨模型迁移的LLM能力。
原文摘要 · Abstract (English)
Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer; off-the-shelf large language models (LLMs) rank the correct disease first in only 35.4% of benchmark cases. Here we show that this expert reasoning can be converted into a scalable AI capability through a governed learning process rather than model training alone. We developed liteOdyssey through Policy Iteration with Human Feedback (PIHF), an in-context policy-learning method adapted from generalized policy iteration in reinforcement learning, in which model failures and expert corrections consolidate into an clinician-gated policy that turns an off-the-shelf LLM into an agentic diagnostic system. We demonstrated that such a policy improved diagnostic accuracy to match the best published systems at a fraction of their deployment footprint, generalized to unseen diseases, transferred across models, and remained under clinician control. Across 1,243 public benchmark cases spanning 722 rare diseases, liteOdyssey ranked the correct disease first in 59.3% of cases versus 26.5% without the policy, with nearly identical gains on the 1,193 cases and 679 diseases excluded from policy development. Ablations showed that gains exceeded automated prompting improvement and source access alone, and the policy transferred without modification across closed- and open-weight models. In 515 Undiagnosed Diseases Network patients, liteOdyssey again improved accuracy, and blinded physicians rated its differentials more often exact and less often unhelpful. Through PIHF, expert reasoning becomes an LLM capability that experts can inspect, revise, and transfer across models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。