arXiv:2410.17035cs.CL2024-10被引 11

用大模型攻击医疗文本脱敏,发现现有方法仍有9%泄露患者信息。

DIRI: Adversarial Patient Reidentification with Large Language Models for Evaluating Clinical Text Anonymization

  • 用大语言模型逆向还原被遮蔽的病历,模拟真实攻击场景。
  • 在三家工具脱敏后,仍成功识别出9%的临床记录对应患者。
  • 适合隐私安全研究者和医疗数据管理者参考。

共享受保护的健康信息(PHI)对推动生物医学研究至关重要。在数据分发前,通常需进行去标识化处理以移除文本中的所有PHI。当前去标识化方法常在高度饱和的数据集上评估,工具表现接近完美,但难以反映真实临床文本的多样性和复杂性,且标注成本高昂,限制了实际应用。为填补这一空白,我们提出一种基于大语言模型(LLM)的对抗性方法,用于重新识别经匿名化处理的临床笔记对应的患者,并引入新的去标识化/再识别(DIRI)评估框架。实验使用来自威利·康奈尔医学院的医疗数据,分别通过规则型Philter及两种深度学习模型——BiLSTM-CRF与ClinicalBERT进行去标识化。尽管ClinicalBERT表现最佳,完全屏蔽了所有个人身份信息(PII),但我们的方法仍成功识别出9%的临床笔记所对应的患者。该研究揭示了现有去标识化技术的重大缺陷,同时提供了一种可迭代改进的评估工具。

原文摘要 · Abstract (English)

Sharing protected health information (PHI) is critical for furthering biomedical research. Before data can be distributed, practitioners often perform deidentification to remove any PHI contained in the text. Contemporary deidentification methods are evaluated on highly saturated datasets (tools achieve near-perfect accuracy) which may not reflect the full variability or complexity of real-world clinical text and annotating them is resource intensive, which is a barrier to real-world applications. To address this gap, we developed an adversarial approach using a large language model (LLM) to re-identify the patient corresponding to a redacted clinical note and evaluated the performance with a novel De-Identification/Re-Identification (DIRI) method. Our method uses a large language model to reidentify the patient corresponding to a redacted clinical note. We demonstrate our method on medical data from Weill Cornell Medicine anonymized with three deidentification tools: rule-based Philter and two deep-learning-based models, BiLSTM-CRF and ClinicalBERT. Although ClinicalBERT was the most effective, masking all identified PII, our tool still reidentified 9% of clinical notes Our study highlights significant weaknesses in current deidentification technologies while providing a tool for iterative development and improvement.

医疗隐私大模型攻击去标识化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。