基于多轴知识与证据验证的中文病历编码框架,提升ICD自动标注准确率与效率。
MKE-Coder: Multi-Axial Knowledge with Evidence Verification in ICD Coding for Chinese EMRs
- 构建四轴疾病知识体系,结合临床证据筛选候选编码。
- 通过掩码语言模型验证知识与证据一致性,减少错误编码。
- 在真实场景中显著提升医生编码速度与准确率,适合医疗AI落地应用。
医学领域自动进行国际疾病分类(ICD)编码的任务已较为成熟,英文语境下取得良好成效,但在处理中文电子病历(EMR)时面临挑战。主要难点在于中文病历书写简洁、结构特殊,难以提取与疾病编码相关的信息;同时,现有方法未能有效利用基于疾病的多轴知识,也缺乏与临床证据的关联。本文提出MKE-Coder:一种面向中文电子病历的多轴知识与证据验证的ICD编码框架。首先,针对诊断识别候选编码,并将其归入四个编码轴下的知识体系;其次,从病历全文中检索对应临床证据,并通过评分模型筛选可信证据;最后,设计基于掩码语言建模的推理模块,验证候选编码所关联的各轴知识是否均得到证据支持,从而给出推荐结果。在多家医院采集的大规模中文病历数据集上进行实验,结果表明MKE-Coder在自动ICD编码任务中表现显著优于现有方法。在模拟真实编码场景的实用性评估中,该方法显著提升了编码员的编码准确率与效率。
原文摘要 · Abstract (English)
The task of automatically coding the International Classification of Diseases (ICD) in the medical field has been well-established and has received much attention. Automatic coding of the ICD in the medical field has been successful in English but faces challenges when dealing with Chinese electronic medical records (EMRs). The first issue lies in the difficulty of extracting disease code-related information from Chinese EMRs, primarily due to the concise writing style and specific internal structure of the EMRs. The second problem is that previous methods have failed to leverage the disease-based multi-axial knowledge and lack of association with the corresponding clinical evidence. This paper introduces a novel framework called MKE-Coder: Multi-axial Knowledge with Evidence verification in ICD coding for Chinese EMRs. Initially, we identify candidate codes for the diagnosis and categorize each of them into knowledge under four coding axes.Subsequently, we retrieve corresponding clinical evidence from the comprehensive content of EMRs and filter credible evidence through a scoring model. Finally, to ensure the validity of the candidate code, we propose an inference module based on the masked language modeling strategy. This module verifies that all the axis knowledge associated with the candidate code is supported by evidence and provides recommendations accordingly. To evaluate the performance of our framework, we conduct experiments using a large-scale Chinese EMR dataset collected from various hospitals. The experimental results demonstrate that MKE-Coder exhibits significant superiority in the task of automatic ICD coding based on Chinese EMRs. In the practical evaluation of our method within simulated real coding scenarios, it has been demonstrated that our approach significantly aids coders in enhancing both their coding accuracy and speed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。