构建首个基于叙事的疾病诊断数据集,提升临床文本自动诊断准确率
MIMIC-SR-ICD11: A Dataset for Narrative-Based Diagnosis
- 提出LL-Rank重排序框架,结合上下文与标签先验优化诊断结果
- 在7种模型上超越基线方法,显著降低标签频率偏差影响
- 适合医疗AI研究者、临床决策支持系统开发者使用
疾病诊断是现代医疗的核心,有助于早期发现急症并及时干预,同时指导生活方式调整和用药以预防或延缓慢性病。自述信息保留了模板化电子健康记录常忽略的临床关键信号,尤其是细微但重要的细节。为推动这一转变,我们引入MIMIC-SR-ICD11,一个基于出院记录构建的大型英文诊断数据集,原生对齐世卫组织ICD-11术语体系。我们进一步提出LL-Rank,一种基于似然的重排序框架,计算每个诊断标签在临床报告上下文中的联合似然(长度归一化),并减去无报告时该标签的先验似然。在七种模型骨干网络上,LL-Rank始终优于强基线生成+映射方法(GenMap)。消融实验表明,其性能提升主要源于基于PMI的评分机制,能有效分离语义兼容性与标签频率偏差。
原文摘要 · Abstract (English)
Disease diagnosis is a central pillar of modern healthcare, enabling early detection and timely intervention for acute conditions while guiding lifestyle adjustments and medication regimens to prevent or slow chronic disease. Self-reports preserve clinically salient signals that templated electronic health record (EHR) documentation often attenuates or omits, especially subtle but consequential details. To operationalize this shift, we introduce MIMIC-SR-ICD11, a large English diagnostic dataset built from EHR discharge notes and natively aligned to WHO ICD-11 terminology. We further present LL-Rank, a likelihood-based re-ranking framework that computes a length-normalized joint likelihood of each label given the clinical report context and subtracts the corresponding report-free prior likelihood for that label. Across seven model backbones, LL-Rank consistently outperforms a strong generation-plus-mapping baseline (GenMap). Ablation experiments show that LL-Rank's gains primarily stem from its PMI-based scoring, which isolates semantic compatibility from label frequency bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。