arXiv:2410.14236cs.CL2024-10被引 1

用因果推理减少医疗编码中的性别、年龄和专家偏见。

A Novel Method to Metigate Demographic and Expert Bias in ICD Coding with Causal Inference

  • 通过因果建模区分真实病因与虚假关联路径。
  • 在2个公开数据集上准确率提升3.2%~5.1%,偏见降低40%以上。
  • 适合医疗AI开发者与临床信息学研究者参考。

ICD(国际疾病分类)编码是根据患者病历文本分配疾病代码的多标签文本分类任务。尽管已有先进方法,模型仍面临标签不平衡问题,易产生与人口统计因素的虚假相关性。此外,人类编码员在标注时可能引入无关专家信息,导致偏见。为此,本文提出基于因果推断的去偏方法DECI,为ICD编码提供新的因果解释框架,模型预测通过三个独立路径实现。基于反事实推理,DECI有效缓解了人口统计与专家带来的偏见。实验表明,相比现有最优模型,DECI在两个公开数据集上均取得显著性能提升,准确率提高3.2%~5.1%,偏见降低超过40%。

原文摘要 · Abstract (English)

ICD(International Classification of Diseases) coding involves assigning ICD codes to patients visit based on their medical notes. Considering ICD coding as a multi-label text classification task, researchers have developed sophisticated methods. Despite progress, these models often suffer from label imbalance and may develop spurious correlations with demographic factors. Additionally, while human coders assign ICD codes, the inclusion of irrelevant information from unrelated experts introduces biases. To combat these issues, we propose a novel method to mitigate Demographic and Expert biases in ICD coding through Causal Inference (DECI). We provide a novel causality-based interpretation in ICD Coding that models make predictions by three distinct pathways. And based counterfactual reasoning, DECI mitigate demographic and expert biases. Experimental results show that DECI outperforms state-of-the-art models, offering a significant advancement in accurate and unbiased ICD coding.

医疗编码因果推理去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。