arXiv:2510.13922cs.LGcs.CL2025-10

让诊断码排序更准,提升医疗编码自动化水平

LTR-ICD: A Ranking-Aware Framework for Automatic ICD Coding

  • 从检索视角建模,同时完成分类与排序任务
  • 主诊断码排序准确率达47%,远超现有方法的20%
  • 适合需要精准排序的临床智能系统使用

临床笔记包含医生在患者就诊时提供的非结构化文本,通常伴随一系列符合国际疾病分类(ICD)的诊断码。正确分配和排序ICD码对医疗诊断与报销至关重要。然而,自动化该任务仍具挑战。现有方法将其视为分类任务,忽略了对不同用途至关重要的代码顺序。本文首次从检索系统视角出发,将此问题重构为分类与排序联合任务。实验表明,所提框架在识别高优先级代码方面表现更优。例如,模型对主诊断码的正确排序准确率为47%,显著高于当前最佳分类器的20%。在分类指标上,微平均和宏平均F1分别为0.6065和0.2904,优于此前最优模型的0.6035和0.2741。

原文摘要 · Abstract (English)

Clinical notes contain unstructured text provided by clinicians during patient encounters. These notes are usually accompanied by a sequence of diagnostic codes following the International Classification of Diseases (ICD). Correctly assigning and ordering ICD codes is essential for medical diagnosis and reimbursement. However, automating this task remains challenging. State-of-the-art methods treated this problem as a classification task, leading to ignoring the order of ICD codes that is essential for different purposes. In this work, as a first attempt, we approach this task from a retrieval system perspective to consider the order of codes, thus formulating this problem as a classification and ranking task. Our results and analysis show that the proposed framework has a superior ability to identify high-priority codes compared to other methods. For instance, our model's accuracy in correctly ranking primary diagnosis codes is 47%, compared to 20% for the state-of-the-art classifier. Additionally, in terms of classification metrics, the proposed model achieves a micro- and macro-F1 scores of 0.6065 and 0.2904, respectively, surpassing the previous best model with scores of 0.6035 and 0.2741.

医疗AI编码系统排序任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。