大模型医疗编码中,隐私保护会降低准确率并加剧性别偏差。
Can large language models be privacy preserving and fair medical coders?
- 在医疗编码任务中引入差分隐私,导致模型性能下降。
- 隐私保护模型在MIMIC-III数据集上微F1分数下降超40%。
- 男性与女性患者间的召回率差距增加超过3%,影响公平性。
在医疗领域部署机器学习算法时,保护患者数据隐私至关重要。差分隐私(DP)是此类场景下的常见隐私保护方法。本文研究了将差分隐私应用于医疗编码(ICD分类)任务中的两个关键权衡。在隐私-效用权衡方面,我们观察到隐私保护模型的性能显著下降,在MIMIC-III数据集上前50个标签的微F1分数降幅超过40%。在隐私-公平性权衡方面,DP模型中男性与女性患者的召回率差距增加了超过3%。深入理解这些权衡,有助于应对真实场景部署中的挑战。
原文摘要 · Abstract (English)
Protecting patient data privacy is a critical concern when deploying machine learning algorithms in healthcare. Differential privacy (DP) is a common method for preserving privacy in such settings and, in this work, we examine two key trade-offs in applying DP to the NLP task of medical coding (ICD classification). Regarding the privacy-utility trade-off, we observe a significant performance drop in the privacy preserving models, with more than a 40% reduction in micro F1 scores on the top 50 labels in the MIMIC-III dataset. From the perspective of the privacy-fairness trade-off, we also observe an increase of over 3% in the recall gap between male and female patients in the DP models. Further understanding these trade-offs will help towards the challenges of real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。