提升罕见病历编码准确率,解决医学标签长尾难题。
CoLa-ICD: A Knowledge-Enhanced Framework for Long-Tail Automated Medical Coding

- 引入外部知识增强标签语义,建模代码间依赖关系
- 在大规模稀疏标签空间中显著提升罕见编码性能
- 适合医疗信息化、临床决策支持系统开发者使用
自动医疗编码将ICD代码分配给临床笔记,但受限于长文档、标签分布不均和术语多样性,尤其对罕见代码(训练样本少、易混淆)挑战更大。本文提出CoLa-ICD,一种知识增强的长尾预测框架:通过引入外部术语丰富标签表示,建模相关代码间的依赖关系,并强化标签语义与临床证据之间的对齐。实验表明,CoLa-ICD在更大、更稀疏的标签空间中取得更优长尾预测效果,在AUC、F1和P@k指标上达到当前最优表现。
原文摘要 · Abstract (English)
Automatic medical coding assigns ICD codes to clinical notes, but it remains challenging due to long documents, imbalanced label distributions, and diverse terms. These challenges are especially severe for rare codes, which have limited training instances and are easily confused with semantically similar labels. We introduce CoLa-ICD, a knowledge-enhanced framework for long-tail prediction. CoLa-ICD enriches ICD labels with external terms, models dependencies among related codes, and learns stronger alignment between label semantics and clinical evidence for long-tail prediction. Experiments show that CoLa-ICD improves long-tail prediction with larger gains in larger and sparser label spaces and achieves state-of-the-art performance in AUC, F1, and P@k. Our code is available at https://github.com/youwillbethebest/Cola-ICD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。