arXiv:2511.14112cs.CLcs.AI2025-11

用合成数据提升罕见病历编码准确率,让长尾代码不再被忽略。

Synthetic Clinical Notes for Rare ICD Codes: A Data-Centric Framework for Long-Tail Medical Coding

  • 基于真实共现模式生成带罕见编码的合成病历
  • 覆盖7902个罕见ICD码,训练集规模扩大数倍
  • 适合医疗数据稀缺场景下的模型公平性优化

从临床文本自动进行ICD编码是医学NLP中的关键任务,但诊断码的极端长尾分布严重制约了性能。如MIMIC-III数据集中,数千个罕见及零样本ICD码严重欠采样,导致宏平均F1分数偏低。本文提出一种数据驱动框架,通过利用真实世界共现模式、ICD描述、同义词、分类体系和相似病历,构建以罕见编码为核心的多标签代码集,并生成90,000条高质量合成出院小结,覆盖7,902个ICD码,显著扩展训练分布。在原始与扩展数据集上微调PLM-ICD和GKI-ICD两个先进Transformer模型,实验表明该方法在保持强微观F1的同时小幅提升宏观F1,优于先前最先进方法。尽管增益相对计算成本较小,结果证明精心设计的合成数据可有效改善长尾ICD编码预测的公平性。

原文摘要 · Abstract (English)

Automatic ICD coding from clinical text is a critical task in medical NLP but remains hindered by the extreme long-tail distribution of diagnostic codes. Thousands of rare and zero-shot ICD codes are severely underrepresented in datasets like MIMIC-III, leading to low macro-F1 scores. In this work, we propose a data-centric framework that generates high-quality synthetic discharge summaries to mitigate this imbalance. Our method constructs realistic multi-label code sets anchored on rare codes by leveraging real-world co-occurrence patterns, ICD descriptions, synonyms, taxonomy, and similar clinical notes. Using these structured prompts, we generate 90,000 synthetic notes covering 7,902 ICD codes, significantly expanding the training distribution. We fine-tune two state-of-the-art transformer-based models, PLM-ICD and GKI-ICD, on both the original and extended datasets. Experiments show that our approach modestly improves macro-F1 while maintaining strong micro-F1, outperforming prior SOTA. While the gain may seem marginal relative to the computational cost, our results demonstrate that carefully crafted synthetic data can enhance equity in long-tail ICD code prediction.

ICD编码合成数据长尾学习医疗NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。