arXiv:2410.01648cs.CL2024-10中稿 · and Presented at: …被引 1

用AI自动隐去病历敏感信息并评估重识别风险,保护患者隐私。

DeIDClinic: A Risk-Aware Pseudonymization Framework for Clinical Text De-identification and Re-identification Risk Assessment

  • 融合BioBERT等模型提升敏感信息识别准确率
  • 在i2b2数据集上多数实体F1超0.96,高风险文档可量化定位
  • 适合医疗、法律等需严控隐私的领域使用

随着敏感文本数据日益增多,亟需可靠的去标识化方法,在保障下游任务可用性的同时实现合规共享。本文提出DeID-Clinic,一个用于临床自由文本自动伪匿名化及重识别风险评估的多层框架。该方法将领域适配的Transformer模型(如BioBERT和ClinicalBERT)集成至MASK去标识化框架,提升对受保护健康信息(PHI)的检测与掩蔽效果。除实体识别外,新增文档级风险评估模块,结合k-匿名、l-多样性、t-接近度、上下文相似性及实体共现分析,量化残余重识别风险。在i2b2 2014去标识化数据集上的实验表明,多个实体类别宏平均F1得分超过0.96,同时可对高风险文档进行定量优先排序以供人工复核。结果表明,神经去标识与显式风险建模结合有效,支持敏感领域的隐私保护数据共享。尽管在临床文本上验证,该框架亦可推广至法律、行政等其他隐私敏感领域。

原文摘要 · Abstract (English)

The increasing availability of sensitive textual data has created an urgent need for robust de-identification methods that enable compliant data sharing while preserving downstream utility. This paper presents DeID-Clinic, a multi-layered framework for automated pseudonymization and re-identification risk assessment of clinical free-text data. Our approach integrates domain-adapted transformer models, including BioBERT and ClinicalBERT, into the MASK de-identification framework to improve the detection and masking of protected health information (PHI). Beyond entity recognition, we introduce a novel document-level risk assessment module that quantifies residual re-identification risk using a combination of k-anonymity, l-diversity, t-closeness, contextual similarity, and entity co-occurrence analysis. Experiments conducted on the i2b2 2014 de-identification dataset demonstrate strong performance, achieving macro-level F1 scores above 0.96 for several entity categories, while enabling quantitative prioritization of high-risk documents for further review. Our results highlight the effectiveness of combining neural de-identification with explicit risk modeling, supporting privacy-preserving data sharing in sensitive domains. Although evaluated on clinical text, the proposed framework is generalizable to other privacy-critical domains such as legal and administrative documents, where reliable pseudonymization and risk-aware anonymization are essential. Keywords{Automated De-Identification, Risk Assessment, Patient Privacy, Pseudonymization, Personal Health Information}

去标识化医疗隐私风险评估NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。