arXiv:2609.09070cs.CLcs.AI2026-09

AI医生在全科诊疗中表现优于人类医生,且能给出更优的诊断和治疗方案。

Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics

  • 用合成病例测试AI、医生与大模型,评估其诊疗全流程表现。
  • AI诊断准确率82.0%,比医生高25个百分点,工作量与治疗评分也更高。
  • 该结果可复现,适合关注临床AI落地的医疗从业者参考。

临床AI评估应涵盖自适应信息获取后的诊断与管理决策。我们在150个合成的波兰语初级护理咨询中,对比了Doctorina、八名医生及四个独立前沿语言模型的表现。Doctorina在首诊一致性上达到82.0%,显著高于医生的57.0%(差异25.0个百分点;95%置信区间17.7–32.7),在主要或参考性鉴别诊断一致性上达97.3%,而医生为85.0%。在149对病例中,标准化诊疗流程评分分别为89.4(AI) vs 66.9(医生),治疗评分83.7 vs 61.2。Doctorina在六组中诊断得分最高,Kimi K3次之;Claude Opus 5则在管理评分中与Opus、Doctorina并列领先。第二次运行验证了医生与人工智能在所有指标上的差距。因此,Doctorina的优势不仅体现在初诊选择,还延伸至自适应咨询后的诊疗流程与初始治疗质量。

原文摘要 · Abstract (English)

Clinical AI evaluation should encompass diagnosis and management after adaptive information gathering. We compared Doctorina, eight physicians and four standalone frontier language models in 150 synthetic Polish-language primary-care consultations. Doctorina achieved 82.0% Top-1 concordance versus 57.0% for physicians (difference, 25.0 percentage points; 95% confidence interval, 17.7-32.7) and 97.3% versus 85.0% primary-or-reference-differential concordance. Across 149 case pairs, normalized workup and treatment scores were 89.4 versus 66.9 and 83.7 versus 61.2. Doctorina had the highest diagnostic point estimates among all six groups; Kimi K3 ranked next, while Claude Opus 5 led the closely spaced management estimates of Opus, Doctorina and Kimi. A second Doctorina execution reproduced the advantages over physicians across all outcomes. Doctorina's advantage over physicians therefore extended from primary-diagnosis selection to higher-rated diagnostic workup and initial treatment after adaptive consultation.

临床AI诊断系统大模型医疗评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。