arXiv:2601.18132cs.AI2026-01

用大模型推理融合技术,提前筛查罕见病风险,提升初诊识别率。

RareAlert: Aligning heterogeneous large language model reasoning for early rare disease risk screening

  • 融合10个大模型的医疗推理,通过机器学习校准加权
  • 在独立测试集上AUC达0.917,优于所有单个大模型和集成方法
  • 适合医院本地部署,兼顾隐私与可扩展性,适用于普遍初诊筛查

罕见病漏诊和延迟诊断仍是重大挑战。初次就诊时医生仅凭有限信息评估风险,若高危患者未被识别,常无法启动针对性检测,导致漏诊。现有初级分诊流程难以可靠发现罕见病患者,亟需普筛机制。本文提出RareAlert,一种基于初诊常规信息预测患者罕见病风险的早期筛查系统。该系统整合10个大语言模型生成的推理,通过机器学习校准并加权这些信号,将对齐后的推理凝练为一个本地可部署的单一模型。为开发与评估,我们构建了真实世界数据集RareBench,包含158,666例病例,覆盖33个Orphanet疾病类别及超7,000种罕见病,涵盖罕见与非罕见表现。结果显示,罕见病识别可重构为对全体患者群体的通用不确定性消解过程。在独立测试集上,基于Qwen3-4B、使用校准推理信号训练的RareAlert达到AUC 0.917,超越最佳机器学习集成及所有评估的大模型(包括GPT-5、DeepSeek-R1、Claude-3.7-Sonnet、o3-mini、Gemini-2.5-Pro、Qwen3-235B)。结果表明大模型医学推理存在多样性,且对齐推理在高度不确定的临床任务中有效。通过将校准推理融入单一模型,RareAlert实现了精准、隐私保护、可扩展的罕见病风险筛查,适用于大规模本地部署。

原文摘要 · Abstract (English)

Missed and delayed diagnosis remains a major challenge in rare disease care. At the initial clinical encounters, physicians assess rare disease risk using only limited information under high uncertainty. When high-risk patients are not recognised at this stage, targeted diagnostic testing is often not initiated, resulting in missed diagnosis. Existing primary care triage processes are structurally insufficient to reliably identify patients with rare diseases at initial clinical presentation and universal screening is needed to reduce diagnostic delay. Here we present RareAlert, an early screening system which predict patient-level rare disease risk from routinely available primary-visit information. RareAlert integrates reasoning generated by ten LLMs, calibrates and weights these signals using machine learning, and distils the aligned reasoning into a single locally deployable model. To develop and evaluate RareAlert, we curated RareBench, a real-world dataset of 158,666 cases covering 33 Orphanet disease categories and more than 7,000 rare conditions, including both rare and non-rare presentations. The results showed that rare disease identification can be reconceptualised as a universal uncertainty resolution process applied to the general patient population. On an independent test set, RareAlert, a Qwen3-4B based model trained with calibrated reasoning signals, achieved an AUC of 0.917, outperforming the best machine learning ensemble and all evaluated LLMs, including GPT-5, DeepSeek-R1, Claude-3.7-Sonnet, o3-mini, Gemini-2.5-Pro, and Qwen3-235B. These findings demonstrate the diversity in LLM medical reasoning and the effectiveness of aligning such reasoning in highly uncertain clinical tasks. By incorporating calibrated reasoning into a single model, RareAlert enables accurate, privacy-preserving, and scalable rare disease risk screening suitable for large-scale local deployment.

罕见病筛查大模型对齐临床决策支持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。