arXiv:2602.08997cs.CYcs.CL2026-02被引 1

旧式去标识化在大模型时代失效,连诊断都能还原患者住址

Paradox of De-identification: A Critique of HIPAA Safe Harbour in the Age of LLMs

  • 用因果图分析病历中隐性身份关联,揭示去标识漏洞
  • 实验证明仅凭诊断信息就能复原患者所在街区
  • 警示医疗隐私保护体系需重构,适合关注数据安全的研究者

隐私是维系医患信任的人类权利。临床笔记记录患者的私密脆弱性与个体特征,用于诊疗协同与研究。根据HIPAA Safe Harbor标准,这些笔记通过移除显式标识符实现去标识化以保护隐私。然而,该标准诞生于以分类表格数据为主的年代,仅关注显式标识符的清除,忽视了身份与准标识符间潜在的统计关联——这类信息可被现代大语言模型捕获。本文首先通过因果图形式化这些关联,随后通过实证手段验证:即使经过清理的病历仍可实现个体重新识别。进一步的诊断消融实验表明,即便移除所有其他信息,模型仍能基于诊断结果预测患者所在社区。本立场论文提出核心问题:当去标识化本质上不完善时,我们如何共同维护医患信任?旨在引发警觉并讨论可行对策。

原文摘要 · Abstract (English)

Privacy is a human right that sustains patient-provider trust. Clinical notes capture a patient's private vulnerability and individuality, which are used for care coordination and research. Under HIPAA Safe Harbor, these notes are de-identified to protect patient privacy. However, Safe Harbor was designed for an era of categorical tabular data, focusing on the removal of explicit identifiers while ignoring the latent information found in correlations between identity and quasi-identifiers, which can be captured by modern LLMs. We first formalize these correlations using a causal graph, then validate it empirically through individual re-identification of patients from scrubbed notes. The paradox of de-identification is further shown through a diagnosis ablation: even when all other information is removed, the model can predict the patient's neighborhood based on diagnosis alone. This position paper raises the question of how we can act as a community to uphold patient-provider trust when de-identification is inherently imperfect. We aim to raise awareness and discuss actionable recommendations.

数据隐私去标识化LLM安全医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。