arXiv:2412.11477cs.LGcs.CL2024-12NeurIPS被引 5

用对比学习提升医疗文本诊断编码准确率

NoteContrast: Contrastive Language-Diagnostic Pretraining for Medical Text

  • 基于对比学习联合预训练诊断代码与病历文本
  • 在MIMIC-III-50等数据集上超越现有最佳模型
  • 适合医疗AI、临床信息处理方向的研究者

准确的医疗记录诊断编码对提升患者护理、医学研究和无误计费至关重要。人工编码耗时且敏感度、特异度低,而病历文本能更精确描述患者状况。因此,自动化诊断编码成为智慧医疗系统的关键。近期长文档Transformer架构使基于注意力的深度学习模型可用于病历判别。此外,对比损失函数已用于带噪声标签的大语言模型与图像模型联合预训练。为进一步提升病历自动判别性能,本文提出一种结合三方面的方法:(i) 基于大规模真实数据集的ICD-10诊断码序列建模,(ii) 医疗文本大语言模型,(iii) 对比预训练以整合诊断码与对应病历文本。实验表明,该对比预训练方法在MIMIC-III-50、MIMIC-III-rare50和MIMIC-III-full诊断编码任务中均优于先前最优模型。

原文摘要 · Abstract (English)

Accurate diagnostic coding of medical notes is crucial for enhancing patient care, medical research, and error-free billing in healthcare organizations. Manual coding is a time-consuming task for providers, and diagnostic codes often exhibit low sensitivity and specificity, whereas the free text in medical notes can be a more precise description of a patients status. Thus, accurate automated diagnostic coding of medical notes has become critical for a learning healthcare system. Recent developments in long-document transformer architectures have enabled attention-based deep-learning models to adjudicate medical notes. In addition, contrastive loss functions have been used to jointly pre-train large language and image models with noisy labels. To further improve the automated adjudication of medical notes, we developed an approach based on i) models for ICD-10 diagnostic code sequences using a large real-world data set, ii) large language models for medical notes, and iii) contrastive pre-training to build an integrated model of both ICD-10 diagnostic codes and corresponding medical text. We demonstrate that a contrastive approach for pre-training improves performance over prior state-of-the-art models for the MIMIC-III-50, MIMIC-III-rare50, and MIMIC-III-full diagnostic coding tasks.

医疗AI诊断编码对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。