比较隐私保护与实用性的荷兰病历去标识化方法
Differentially Private De-identification of Dutch Clinical Notes: A Comparative Evaluation

- 对比差分隐私、命名实体识别和大模型在病历去标识中的表现
- 大模型预处理后加差分隐私,隐私与实用性平衡最佳
- 适合医疗数据合规研究者及隐私保护技术开发者
在符合GDPR和HIPAA等法规的前提下,保护临床文本中的患者隐私对数据二次使用至关重要。尽管人工去标识仍是金标准,但成本高、效率低,推动了自动化方法的发展。现有自动化方案多基于命名实体识别(NER)识别需脱敏信息。差分隐私(DP)提供形式化隐私保障,而大语言模型(LLMs)在临床去标识中日益广泛应用。本文首次对荷兰语临床文本的差分隐私、NER和大模型方法进行对比研究,分别评估其性能,并探索在差分隐私前使用NER或大模型预处理的混合策略。评估指标包括隐私泄露程度和外部任务表现(实体与关系分类)。结果表明,仅使用差分隐私会显著降低数据实用性,但结合语言学预处理,尤其是基于大模型的去标识,能显著改善隐私-实用性的权衡。
原文摘要 · Abstract (English)
Protecting patient privacy in clinical narratives is essential for enabling secondary use of healthcare data under regulations such as GDPR and HIPAA. While manual de-identification remains the gold standard, it is costly and slow, motivating the need for automated methods that combine privacy guarantees with high utility. Most automated text de-identification pipelines employed named entity recognition (NER) to identify protected entities for redaction. Although methods based on differential privacy (DP) provide formal privacy guarantees, more recently also large language models (LLMs) are increasingly used for text de-identification in the clinical domain. In this work, we present the first comparative study of DP, NER, and LLMs for Dutch clinical text de-identification. We investigate these methods separately as well as hybrid strategies that apply NER or LLM preprocessing prior to DP, and assess performance in terms of privacy leakage and extrinsic evaluation (entity and relation classification). We show that DP mechanisms alone degrade utility substantially, but combining them with linguistic preprocessing, especially LLM-based redaction, significantly improves the privacy-utility trade-off.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。