专为罕见病诊断设计的智能模型,提升准确率并减少误诊。
A Specialized Large Language Model for Clinical Reasoning and Diagnosis in Rare Diseases
- 基于临床推理数据与分阶段训练,融合图谱检索增强诊断。
- 在多中心病历中达到顶尖准确率,噪声环境下仍稳定表现。
- 可解释性强,辅助医生识别非表型关键证据,适合临床辅助使用。
罕见病影响全球数亿人,但诊断常需数年。传统流程将证据提取与推理诊断分离,通用或医学大模型受限于真实电子健康记录稀缺、知识过时及幻觉问题。我们构建大规模领域专用临床语料库与医师验证的推理数据集,通过分阶段指令微调、思维链学习与图谱引导检索,开发出RareSeek R1模型。在多中心电子病历文本与公开基准测试中,该模型实现最先进的准确率、强泛化能力与在噪声或重叠表型下的稳定性。当病历与优先变异基因结合时,增强检索带来最大收益,有效消除歧义并使候选诊断与致病机制对齐。人类评估显示其性能媲美经验丰富的医生,并在辅助使用中持续提升。值得注意的是,透明推理揭示了非表型证据(中位占比23.1%,如影像、干预措施、功能检测)对多数正确诊断的关键作用。本工作推动以叙事为中心、知识融合的推理范式,缩短诊断旅程,实现可审计、可临床落地的决策支持。
原文摘要 · Abstract (English)
Rare diseases affect hundreds of millions worldwide, yet diagnosis often spans years. Convectional pipelines decouple noisy evidence extraction from downstream inferential diagnosis, and general/medical large language models (LLMs) face scarce real world electronic health records (EHRs), stale domain knowledge, and hallucinations. We assemble a large, domain specialized clinical corpus and a clinician validated reasoning set, and develop RareSeek R1 via staged instruction tuning, chain of thought learning, and graph grounded retrieval. Across multicenter EHR narratives and public benchmarks, RareSeek R1 attains state of the art accuracy, robust generalization, and stability under noisy or overlapping phenotypes. Augmented retrieval yields the largest gains when narratives pair with prioritized variants by resolving ambiguity and aligning candidates to mechanisms. Human studies show performance on par with experienced physicians and consistent gains in assistive use. Notably, transparent reasoning highlights decisive non phenotypic evidence (median 23.1%, such as imaging, interventions, functional tests) underpinning many correct diagnoses. This work advances a narrative first, knowledge integrated reasoning paradigm that shortens the diagnostic odyssey and enables auditable, clinically translatable decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。