用多位专家模型共识提升医疗AI决策能力,更准更稳。
Second Opinion Matters: Towards Adaptive Clinical AI via the Consensus of Expert Model Ensemble
- 组建多个专业医疗模型,像医生会诊一样协同决策。
- 在多个医学测试中准确率显著超越单模型,最高提9.1%。
- 适合对可靠性要求高的医疗AI系统,尤其擅长诊断推理。
尽管大语言模型在临床中应用日益广泛,现有方法仍严重依赖单一模型架构。为应对模型过时和僵化依赖的风险,我们提出一种新型框架——共识机制。该机制模拟临床分诊与多学科会诊,通过一组专业化医疗专家代理构成集成系统,提升临床决策质量并保持强适应性。此架构可仅通过内部模型配置优化成本、延迟或性能。我们在三个医学评估基准(MedMCQA、MedQA、MedXpertQA Text)和一个鉴别诊断数据集DDX+上进行严格评估。在MedXpertQA上,共识机制准确率达61.0%,优于OpenAI的O3(53.5%)和Google的Gemini 2.5 Pro(45.9%)。在MedQA和MedMCQA上分别提升3.4%和9.1%。诊断生成方面,其召回率与精确率也更高(F1$_\mathrm{consensus}$ = 0.326 vs. F1$_\mathrm{O3-high}$ = 0.2886),DDX top-1准确率达52.0%,高于O3-high的45.2%。
原文摘要 · Abstract (English)
Despite the growing clinical adoption of large language models (LLMs), current approaches heavily rely on single model architectures. To overcome risks of obsolescence and rigid dependence on single model systems, we present a novel framework, termed the Consensus Mechanism. Mimicking clinical triage and multidisciplinary clinical decision-making, the Consensus Mechanism implements an ensemble of specialized medical expert agents enabling improved clinical decision making while maintaining robust adaptability. This architecture enables the Consensus Mechanism to be optimized for cost, latency, or performance, purely based on its interior model configuration. To rigorously evaluate the Consensus Mechanism, we employed three medical evaluation benchmarks: MedMCQA, MedQA, and MedXpertQA Text, and the differential diagnosis dataset, DDX+. On MedXpertQA, the Consensus Mechanism achieved an accuracy of 61.0% compared to 53.5% and 45.9% for OpenAI's O3 and Google's Gemini 2.5 Pro. Improvement was consistent across benchmarks with an increase in accuracy on MedQA ($Δ\mathrm{Accuracy}_{\mathrm{consensus\text{-}O3}} = 3.4\%$) and MedMCQA ($Δ\mathrm{Accuracy}_{\mathrm{consensus\text{-}O3}} = 9.1\%$). These accuracy gains extended to differential diagnosis generation, where our system demonstrated improved recall and precision (F1$_\mathrm{consensus}$ = 0.326 vs. F1$_{\mathrm{O3\text{-}high}}$ = 0.2886) and a higher top-1 accuracy for DDX (Top1$_\mathrm{consensus}$ = 52.0% vs. Top1$_{\mathrm{O3\text{-}high}}$ = 45.2%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。