用迭代提问模拟真实问诊,让AI诊断准确率提升至80%。
Sequential Diagnosis with Language Models
- 构建逐步提问的诊断评测基准,模拟医生实时推理过程。
- 搭配专用调度器后,诊断准确率达80%,成本降低70%。
- 适用于各类大模型,推动临床诊断更精准、更经济。
人工智能有望扩大专家医疗知识的可及性。然而,现有语言模型评估多依赖静态病例和选择题,难以反映真实临床中循证医学的复杂性。在实际诊疗中,医生会逐步提出并修正诊断假设,根据新信息调整后续问题与检查,并综合动态证据做出最终判断。为此,我们构建了序列化诊断基准,将304例《新英格兰医学杂志》临床病理讨论案例转化为分步诊断流程。医生或AI从简要摘要开始,需逐项向一个守门模型请求信息,仅当明确询问时才揭示结果。评估不仅关注诊断准确率,还考量就诊与检测成本。我们提出MAI诊断调度器(MAI-DxO),一种模型无关的调度系统,能模拟多医生协作,生成可能鉴别诊断,并策略性选择高价值、低成本检测。当与OpenAI的o3模型结合,其诊断准确率达80%——是普通医生平均20%的四倍。同时,相比医生减少20%诊断成本,相比未优化的o3模型降低70%。在追求最高精度时,准确率可达85.5%。该效果在OpenAI、Gemini、Claude、Grok、DeepSeek及Llama系列模型上均具泛化能力。研究证明,引导AI进行迭代思考与审慎行动,可显著提升诊断精度与成本效益。
原文摘要 · Abstract (English)
Artificial intelligence holds great promise for expanding access to expert medical knowledge and reasoning. However, most evaluations of language models rely on static vignettes and multiple-choice questions that fail to reflect the complexity and nuance of evidence-based medicine in real-world settings. In clinical practice, physicians iteratively formulate and revise diagnostic hypotheses, adapting each subsequent question and test to what they've just learned, and weigh the evolving evidence before committing to a final diagnosis. To emulate this iterative process, we introduce the Sequential Diagnosis Benchmark, which transforms 304 diagnostically challenging New England Journal of Medicine clinicopathological conference (NEJM-CPC) cases into stepwise diagnostic encounters. A physician or AI begins with a short case abstract and must iteratively request additional details from a gatekeeper model that reveals findings only when explicitly queried. Performance is assessed not just by diagnostic accuracy but also by the cost of physician visits and tests performed. We also present the MAI Diagnostic Orchestrator (MAI-DxO), a model-agnostic orchestrator that simulates a panel of physicians, proposes likely differential diagnoses and strategically selects high-value, cost-effective tests. When paired with OpenAI's o3 model, MAI-DxO achieves 80% diagnostic accuracy--four times higher than the 20% average of generalist physicians. MAI-DxO also reduces diagnostic costs by 20% compared to physicians, and 70% compared to off-the-shelf o3. When configured for maximum accuracy, MAI-DxO achieves 85.5% accuracy. These performance gains with MAI-DxO generalize across models from the OpenAI, Gemini, Claude, Grok, DeepSeek, and Llama families. We highlight how AI systems, when guided to think iteratively and act judiciously, can advance diagnostic precision and cost-effectiveness in clinical care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。