arXiv:2607.23290cs.AI2026-07

用多个模型的分歧推理,提升罕见病诊疗的准确性与可靠性。

RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning

  • 通过整合多个大模型的差异性推理路径,实现罕见病全流程决策。
  • 在15万例真实病例上,筛查准确率AUC达0.917,诊断与治疗准确率分别达65.5%和89.8%。
  • 适合临床决策支持系统研发者及罕见病诊疗研究者参考。

罕见病临床决策面临巨大挑战,因表现异质、证据稀少、专家稀缺,全程存在持续不确定性。现有AI系统多聚焦孤立任务,依赖后续检查而非初始信息。本文提出RareLens,不靠单一模型规模扩展,而是利用多个大模型间差异化的推理路径及其互补的错误模式。基于涵盖33类孤儿病、超7000种疾病的157,525例真实病例数据集RarelensBench,RareLens在风险筛查、诊断、治疗规划和预后预测四个阶段均超越包括GPT-5、DeepSeek-R1、Claude-3.7-Sonnet和Gemini-2.5-Pro在内的所有前沿模型。其筛查阶段AUC达0.917,诊断与治疗的top-1准确率分别为65.5%和89.8%。外部评估中,1,287例病例与23名医生参与,自主运行的RareLens及辅助医生的版本均优于未受助医生,表明高效人机协作需超越简单输出提供。研究证明,多样化模型推理是可挖掘的信息源,为高不确定性临床环境下的可靠AI系统提供通用策略。

原文摘要 · Abstract (English)

Rare diseases represent one of the most challenging settings for clinical decision-making, where heterogeneous presentations, sparse evidence and limited expertise create persistent uncertainty throughout the care pathway. Although artificial intelligence could help, existing systems largely address isolated tasks, particularly diagnosis, and usually rely on downstream investigations rather than information available at initial presentation. Here we show that clinical AI performance under uncertainty can be improved not by scaling a single model, but by exploiting the diversity of multiple imperfect reasoning systems. Across heterogeneous large language models, we identify divergent reasoning trajectories with complementary error patterns and develop RareLens, which learns to reconcile these perspectives into actionable decisions across four stages of rare disease care: risk screening, diagnosis, treatment planning and prognosis prediction. Built on RarelensBench, a real-world dataset of 157,525 cases spanning all 33 Orphanet categories and more than 7,000 conditions, RareLens outperformed every frontier model tested, including GPT-5, DeepSeek-R1, Claude-3.7-Sonnet and Gemini-2.5-Pro, across all stages. It achieved an area under the curve of 0.917 for screening and top-1 accuracies of 65.5% and 89.8% for diagnosis and treatment. In an external evaluation involving 1,287 cases and 23 physicians, autonomous RareLens and physicians assisted by RareLens both outperformed unaided physicians, while demonstrating that effective human-AI collaboration requires more than simply providing model outputs. These findings establish divergent model reasoning as an exploitable source of information and suggest a general strategy for building AI systems that operate reliably under high clinical uncertainty.

罕见病大模型临床决策人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。