arXiv:2504.21035cs.CRcs.CL2025-04被引 29

现有文本去敏方法看似安全,实则易被重构身份,威胁隐私。

A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage

  • 构建新评估框架,检测文本中隐性语义线索导致的重识别风险。
  • 微软Azure工具在MedQA数据集上仅保护26%敏感信息(泄露74%)。
  • 适合关注隐私安全、数据发布合规的研究者与从业者。

文本数据去敏通常通过移除个人身份信息(PII)或生成合成数据来实现,但其有效性常仅通过显式标识符泄露程度衡量,忽视了可能导致重识别的细微文本特征。本文提出新评估框架,量化数据发布后的个体隐私风险。结果表明,看似无害的日常社交行为等辅助信息,足以推断出年龄、药物使用史等敏感属性。例如,微软Azure的商业级PII清除工具在MedQA数据集上未能保护74%的信息。尽管差分隐私可部分缓解风险,但严重损害下游任务的文本可用性。研究揭示当前去敏技术带来虚假的安全感,亟需防范语义层面信息泄露的更强方法。

原文摘要 · Abstract (English)

Sanitizing sensitive text data typically involves removing personally identifiable information (PII) or generating synthetic data under the assumption that these methods adequately protect privacy; however, their effectiveness is often only assessed by measuring the leakage of explicit identifiers but ignoring nuanced textual markers that can lead to re-identification. We challenge the above illusion of privacy by proposing a new framework that evaluates re-identification attacks to quantify individual privacy risks upon data release. Our approach shows that seemingly innocuous auxiliary information -- such as routine social activities -- can be used to infer sensitive attributes like age or substance use history from sanitized data. For instance, we demonstrate that Azure's commercial PII removal tool fails to protect 74\% of information in the MedQA dataset. Although differential privacy mitigates these risks to some extent, it significantly reduces the utility of the sanitized text for downstream tasks. Our findings indicate that current sanitization techniques offer a \textit{false sense of privacy}, highlighting the need for more robust methods that protect against semantic-level information leakage.

隐私保护数据去敏重识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。