用安全导向框架提升AI诊断的可靠性,防止漏诊致命疾病。
A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis

- 分角色协作的LLM系统,通过证据检索与验证关卡生成诊断。
- 在急诊病例中,前3名诊断准确率达85.7%,漏诊高危病种率降低至22.0%。
- 适合需要高安全性的临床决策支持,尤其急诊和重症场景。
诊断错误是患者安全的重大威胁,而现有大语言模型(LLM)系统常将诊断视为一次性预测任务,缺乏对高风险遗漏病种的防护及推理过程的严格验证。本文提出AegisDx——一种面向安全的假说演绎式临床推理框架。该框架通过角色化合约、结构化中间输出、证据检索接口与验证门控,协调多个专用LLM组件,实现全面的鉴别诊断生成,强制筛查高危“必须排除”病症,依据真实医学证据验证推理,并规划可操作的下一步行动。我们在三个层面评估:基于NEJM和JAMA文献案例,以GPT-oss-120B为共享基底,AegisDx在JAMA病例中Top-3准确率为59.9%(单体模型52.1%),在NEJM病例中为62.7%(单体51.4%);在《急诊医学年鉴》案例中,准确率达85.7%(单体68.6%)。在对抗医师共识的“必须排除”病种集合时,AegisDx在78.0%的案例中至少覆盖一个高危病种(单体52.0%)。在对耶鲁纽黑文医疗系统43份真实急诊科病历的盲评中,相较GPT-5,AegisDx将医师评分的综合安全得分从4.31提升至4.55(5分制,调整后p=2.1×10⁻⁴),定性上显著提升高危病种识别与推理安全性。结果表明,将诊断AI设计为以安全为核心的推理框架,而非仅优化预测准确率,可为急症临床流程提供更安全、透明且具临床意义的床旁辅助支持。
原文摘要 · Abstract (English)
Diagnostic error is a major threat to patient safety, yet current large language model (LLM) systems often treat diagnosis as a one-shot prediction task, lacking safeguards against missed high-risk alternatives or rigorous verification of their reasoning. Here, we present AegisDx, a safety-oriented framework for hypothetico-deductive clinical reasoning. AegisDx coordinates specialized LLM components through role-specific contracts, structured intermediate outputs, evidence-retrieval interfaces, and verification gates to generate broad differential diagnoses, enforce explicit screening for dangerous "must-not-miss" conditions, verify reasoning against grounded medical evidence, and structure actionable next steps. We evaluated AegisDx across three layers. On literature-derived case reports from NEJM and JAMA, with GPT-oss-120B as the shared backbone, Top-3 diagnostic accuracy was 59.9% versus 52.1% for the standalone LLM on JAMA cases and 62.7% versus 51.4% on NEJM cases. On cases from Annals of Emergency Medicine, Top-3 accuracy was 85.7% versus 68.6%; against physician-consensus must-not-miss diagnosis sets, AegisDx captured at least one such condition among its top three diagnoses in 78.0% of cases versus 52.0%. In a blinded physician evaluation of 43 real-world emergency department notes from the Yale New Haven Health System compared against GPT-5, AegisDx improved the physician-rated composite safety score from 4.31 to 4.55 on a 5-point scale (adjusted p = 2.1x10^-4), with qualitative gains in must-not-miss identification and reasoning safety. Our findings suggest that engineering diagnostic AI as a safety-oriented reasoning framework, rather than optimizing raw predictive accuracy alone, can provide a safer, more transparent, and clinically meaningful layer of bedside decision support for acute care workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。