arXiv:2607.28788cs.AI2026-07

构建首个基于急诊入院证据的开放式诊断基准,评估AI在有限信息下真实推理能力。

EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses

论文配图:EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses
图 1 · 摘自论文原文
  • 以入院时可得记录为依据,用急诊期间诊断而非出院诊断作为监督信号
  • 现有模型仅能正确推断3-31%需推理的诊断,最佳模型推理召回率56%
  • 适合研究临床AI推理、医疗大模型评估与急诊决策支持系统的开发者

医院急诊入院时的临床诊断必须在证据有限且不完整的情况下快速做出。现有诊断预测基准难以适配此场景:它们限制预测为封闭代码集,排除自由文本病历,且以包含全部住院过程的出院诊断作为监督标签。我们提出EarlyDx,一个基于MIMIC-IV数据集的大型开放式早期诊断基准,涵盖154,834例急诊就诊记录。每例仅使用入院时刻$t_0$前的记录,并以急诊期间记录的诊断作为监督标签,而非出院诊断。通过大模型审计器对每个自由文本标签进行验证,判断其是否被证据完全支持、部分支持或无支持;评估仅计入完全支持的标签。在语义级大模型评判协议下,所有测试系统——包括前沿通用、医学专用及领域微调模型——均无法可靠地从入院证据中生成诊断。零样本模型主要依赖提取,仅能恢复3-31%需推理的诊断;微调后推理召回率提升至56%,但仍存显著差距;在时间敏感疾病上,无一系统达到临床医生的敏感性与精确性平衡。我们已公开完整构建与评估流程。

原文摘要 · Abstract (English)

Clinical diagnosis at hospital admission must be made rapidly from limited, incomplete evidence. Existing diagnosis-prediction benchmarks are poorly suited to this setting: they restrict prediction to closed code sets, exclude free-text notes, and supervise with discharge diagnoses that incorporate the full inpatient course. We introduce EarlyDx, a large-scale benchmark for open-ended early diagnosis, built from 154,834 emergency department encounters in MIMIC-IV. Each encounter is restricted to records available at admission time $t_0$ and supervised by the diagnoses recorded during the ED encounter rather than at discharge. An LLM auditor further verifies every free-text label as supported, partially supported, or unsupported by that evidence; the primary evaluation scores only fully supported labels. Under a semantic LLM-as-judge protocol, no evaluated system --- frontier general, medical-specialized, or in-domain post-trained --- synthesizes admission-time evidence reliably. Zero-shot models score largely by extraction, recovering only 3-31% of diagnoses that must be inferred rather than read from the record; post-training raises inference-dependent recall to 56%, but a sizeable margin remains, and on time-critical conditions no system attains a clinician's balance of sensitivity and precision. We release the full construction and evaluation pipeline at here.

诊断生成医疗大模型急诊决策评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。