arXiv:2608.00935cs.LG2026-08

将疾病编码转化为可解释的低维表示,兼顾预测性能与临床可读性。

xMICD: Explainable Representation of Multiple ICD Codes

论文配图:xMICD: Explainable Representation of Multiple ICD Codes
图 1 · 摘自论文原文
  • 融合诊断分组与编码嵌入相似性,动态分配代码到临床组别
  • 在多个任务上达到与ICD2Vec相当的预测性能
  • 每维对应可识别的诊断组,适合医疗场景可解释性需求

电子健康记录(EHR)广泛用于机器学习驱动的临床风险预测。国际疾病分类(ICD)编码提供了患者诊断的结构化信息,但其有效表示仍具挑战。现有方法在预测性能与可解释性之间存在权衡:基于分组的表示虽可解释但易丢失信息,基于嵌入的表示预测能力强但难以解读。本文提出可解释的多ICD编码表示方法(xMICD),从一组ICD编码构建低维患者表征。xMICD结合临床有意义的诊断分组与预训练ICD嵌入空间中的相似性,通过相似性赋权而非二值归属,将编码分配至分组,生成反映患者诊断与临床组别契合度的特征。在大规模EHR数据集上的实验表明,xMICD在多个临床预测任务中达到与ICD2Vec等嵌入方法相当的预测性能,同时保留了临床可解释性——每个维度对应一个可识别的诊断组。该方法为将嵌入语义关系融入可解释临床特征空间提供了实用路径。

原文摘要 · Abstract (English)

Electronic Health Records (EHRs) are widely used for clinical risk prediction using machine learning. International Classification of Diseases (ICD) codes provide structured information about patient diagnoses, but representing them effectively remains challenging. Existing approaches often face a trade-off between predictive performance and interpretability: grouping-based representations are interpretable but may lose information, while embedding-based representations achieve strong predictive performance but are difficult to interpret. We propose Explainable Representation of Multiple ICD Codes (xMICD), a method for constructing low-dimensional patient representations from sets of ICD codes. xMICD combines clinically meaningful diagnostic groupings with similarity in a pre-trained ICD embedding space. Instead of using binary group membership, the method assigns codes to groups via similarity-based relative assignments, yielding features that reflect how closely a patient's diagnoses align with each clinical group. Experiments on large-scale EHR datasets demonstrate that xMICD achieves predictive performance comparable to embedding-based representations such as ICD2Vec across multiple clinical prediction tasks. At the same time, the resulting features remain clinically interpretable because each dimension corresponds to a recognizable diagnostic group. xMICD therefore provides a practical way to integrate embedding-based semantic relationships into interpretable clinical feature spaces for machine learning models.

医疗AI可解释性疾病编码表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。