用患者知识图谱提升疾病编码可解释性,效果更好且更省力。
Structured Information Matters: Explainable ICD Coding with Patient-Level Knowledge Graphs
- 构建患者级知识图谱,压缩原文23%体积保留90%信息。
- 在ICD-9编码任务中,宏平均F1最高提升3.20%,训练更快。
- 图谱结构增强模型可解释性,适合医疗AI可解释性研究者。
将临床文档映射到标准化临床术语是重要任务,有助于临床研究、医院管理与患者护理。但人工编码耗时费力,难以规模化。自动编码可缓解此问题,但需处理高维、长尾的ICD编码空间。现有方法多依赖外部知识增强输出表示,而对输入文档的结构化建模较少。本文通过构建患者级知识图谱(KG),以结构化方式表征病历,仅用原始文本23%的体量保留90%信息。将该图谱集成至PLM-ICD架构,在主流基准上实现最高3.20%的宏平均F1提升,并提高训练效率。结果归因于图谱中实体与关系的多样性,同时显著增强了模型可解释性。
原文摘要 · Abstract (English)
Mapping clinical documents to standardised clinical vocabularies is an important task, as it provides structured data for information retrieval and analysis, which is essential to clinical research, hospital administration and improving patient care. However, manual coding is both difficult and time-consuming, making it impractical at scale. Automated coding can potentially alleviate this burden, improving the availability and accuracy of structured clinical data. The task is difficult to automate, as it requires mapping to high-dimensional and long-tailed target spaces, such as the International Classification of Diseases (ICD). While external knowledge sources have been readily utilised to enhance output code representation, the use of external resources for representing the input documents has been underexplored. In this work, we compute a structured representation of the input documents, making use of document-level knowledge graphs (KGs) that provide a comprehensive structured view of a patient's condition. The resulting knowledge graph efficiently represents the patient-centred input documents with 23\% of the original text while retaining 90\% of the information. We assess the effectiveness of this graph for automated ICD-9 coding by integrating it into the state-of-the-art ICD coding architecture PLM-ICD. Our experiments yield improved Macro-F1 scores by up to 3.20\% on popular benchmarks, while improving training efficiency. We attribute this improvement to different types of entities and relationships in the KG, and demonstrate the improved explainability potential of the approach over the text-only baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。