用180万患者数据训练医疗编码模型,提升诊断代码自动提取准确率。
A medical coding language model trained on clinical narratives from a population-wide cohort of 1.8 million patients
- 基于580万病历训练大语言模型,融合临床记录与检验结果预测编码
- 微F1达71.8%,前十候选召回率达95.5%,部分专科表现优异
- 发现二级诊断普遍漏码,模型可识别数千例未编码病例
医疗编码将临床文档转化为用于计费、研究和公共卫生的标准化代码,但人工编码耗时且易出错。现有自动化方法依赖小规模数据集,难以反映真实患者异质性。我们基于丹麦东部近全科别、180万患者(2006–2016年)的580万份电子健康记录训练了一个语言模型,从临床笔记、药物和实验室结果中预测ICD-10代码。在27万例保留测试患者上,模型达到微F1 71.8%和前10名召回率95.5%。性能按专科差异显著(F1:53–91%),诊断标准明确的专科表现更优。以次要诊断为主的代码F1明显偏低。针对自杀相关行为、体重障碍和高血压三类代码,模型识别出数千例未编码病例,其中76–86%经人工核查确认有效,表明系统性漏码而非模型错误。这提示该时期丹麦东部存在二级诊断编码不足,可能影响流行病学研究、公共卫生监测及共病理解。类似时间约束与报销结构在其他医疗系统中普遍存在,暗示此问题或具普遍性。该模型可自动处理约50%病例,并为其余提供高精度建议,或能有效捕捉被遗漏的次要疾病。
原文摘要 · Abstract (English)
Medical coding translates clinical documentation into standardized codes for billing, research, and public health, but manual coding is time-consuming and error-prone. Existing automation efforts rely on small datasets that poorly represent real-world patient heterogeneity. We trained a language model on 5.8 million electronic health records from 1.8 million patients across nearly all specialties in Eastern Denmark (2006--2016) to predict ICD-10 codes from clinical notes, medications, and laboratory results. Evaluated on 270,000 held-out patients, the model achieved a micro F1 of 71.8% and a top-10 recall of 95.5%. Performance varied by specialty (F1: 53--91%), with higher scores in specialties with well-defined diagnostic criteria. Codes appearing predominantly as secondary diagnoses had markedly lower F1 scores. For three such codes (suicide-related behaviors, weight disorders, and hypertension), the model identified thousands of uncoded cases, of which 76-86% were confirmed valid upon manual review, suggesting systematic under-coding rather than model error. These findings suggest under-coding of secondary diagnoses in Eastern Denmark during this period, with potential implications for epidemiological research, public health surveillance, and understanding of multimorbidity. Similar time constraints and reimbursement structures in other healthcare systems suggest this may not be isolated to this dataset. The model can automate coding for approximately 50% of cases and provide accurate suggestions for most others, and may offer a practical solution to help capture missed secondary conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。