arXiv:2511.09559cs.LGcs.AI2025-11

通过概率引导图注意力提升罕见疾病编码准确率

Enhancing Rare Codes via Probability-Biased Directed Graph Attention for Long-Tail ICD Coding

  • 构建概率加权有向图,仅允许常见码向罕见码传递信息
  • 在三个数据集上大幅提高罕见码的宏平均F1分数
  • 结合大模型生成的语义描述,增强临床上下文理解

自动化国际疾病分类(ICD)编码旨在为临床文档分配多个疾病代码,在医疗信息学中至关重要。然而,其性能受制于ICD本体的极端长尾分布:少数常见代码占据主导,而数千个罕见代码样本极少。为此,我们提出概率偏置有向图注意力模型(ProBias),将代码分为常见与罕见两类,并仅允许信息从常见码流向罕见码。边权重由条件共现概率决定,引导注意力机制将临床相关信号注入罕见码表示。为进一步提升输入语义质量,我们利用大语言模型为ICD代码生成丰富文本描述,补充统计共现信号。在自动ICD编码任务中,该方法显著改善了罕见码的表征与预测效果,在三个基准数据集上达到当前最优性能,尤其在宏观平均F1分数上取得显著提升。

原文摘要 · Abstract (English)

Automated international classification of diseases (ICD) coding aims to assign multiple disease codes to clinical documents and plays a critical role in healthcare informatics. However, its performance is hindered by the extreme long-tail distribution of the ICD ontology, where a few common codes dominate while thousands of rare codes have very few examples. To address this issue, we propose a Probability-Biased Directed Graph Attention model (ProBias) that partitions codes into common and rare sets and allows information to flow only from common to rare codes. Edge weights are determined by conditional co-occurrence probabilities, which guide the attention mechanism to enrich rare-code representations with clinically related signals. To provide higher-quality semantic representations as model inputs, we further employ large language models to generate enriched textual descriptions for ICD codes, offering external clinical context that complements statistical co-occurrence signals. Applied to automated ICD coding, our approach significantly improves the representation and prediction of rare codes, achieving state-of-the-art performance on three benchmark datasets. In particular, we observe substantial gains in macro-averaged F1 score, a key metric for long-tail classification.

ICD编码长尾学习图注意力大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。