arXiv:2502.13342cs.CL2025-02被引 3

提出九类间接标识符分类,提升医疗文本去标识化能力

Beyond De-Identification: A Structured Approach for Defining and Detecting Indirect Identifiers in Medical Texts

  • 构建九类间接标识符分类体系,覆盖亲友、医护人员等潜在攻击者
  • 标注100份MIMIC-III出院记录,共6199个标识符片段
  • 开源标注指南与数据,助力隐私保护研究

为科学共享敏感医疗文本,需有效保护患者与医护人员隐私。文本去标识化尤其困难,因存在多样化的非结构化直接与间接标识符。为降低再识别风险,本文提出涵盖九类间接标识符的分类体系,以应对包括熟人、家属及医疗人员在内的不同潜在攻击者。基于该体系,我们对100份MIMIC-III出院记录进行标注,共生成6,199个标注片段,并提出用于识别间接标识符的基线模型。相关标注指南、标注范围(含6,199个标注)及对应MIMIC-III文档编号将公开发布,以支持该领域的后续研究。

原文摘要 · Abstract (English)

Sharing sensitive texts for scientific purposes requires appropriate techniques to protect the privacy of patients and healthcare personnel. Anonymizing textual data is particularly challenging due to the presence of diverse unstructured direct and indirect identifiers. To mitigate the risk of re-identification, this work introduces a schema of nine categories of indirect identifiers designed to account for different potential adversaries, including acquaintances, family members and medical staff. Using this schema, we annotate 100 MIMIC-III discharge summaries and propose baseline models for identifying indirect identifiers. We will release the annotation guidelines, annotation spans (6,199 annotations in total) and the corresponding MIMIC-III document IDs to support further research in this area.

医疗文本隐私保护去标识化间接标识符

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。