梳理文本隐私度量方法,帮研究者选对评估标准。
How do we measure privacy in text? A survey of text anonymization metrics
- 系统分析47篇论文的6类隐私概念及其度量方式
- 发现现有方法与HIPAA/GDPR法律标准存在差距
- 为学术界提供可比、合规的隐私评估指南
本文通过系统性综述,澄清并整合文本隐私保护评估指标。尽管文本匿名化对敏感数据领域的自然语言处理研究至关重要,但如何评估其隐私保护效果仍是开放问题。我们手动审查了47篇报告隐私度量的论文,识别并比较了六种不同的隐私概念,分析其相关指标如何捕捉不同层面的隐私风险。进一步评估这些概念与法律标准(HIPAA和GDPR)以及基于人机交互研究的用户期望之间的契合度。研究为理解隐私评估方法提供了实践指导,揭示了当前实践中的不足,旨在推动更稳健、可比且符合法律要求的文本匿名化评估体系。
原文摘要 · Abstract (English)
In this work, we aim to clarify and reconcile metrics for evaluating privacy protection in text through a systematic survey. Although text anonymization is essential for enabling NLP research and model development in domains with sensitive data, evaluating whether anonymization methods sufficiently protect privacy remains an open challenge. In manually reviewing 47 papers that report privacy metrics, we identify and compare six distinct privacy notions, and analyze how the associated metrics capture different aspects of privacy risk. We then assess how well these notions align with legal privacy standards (HIPAA and GDPR), as well as user-centered expectations grounded in HCI studies. Our analysis offers practical guidance on navigating the landscape of privacy evaluation approaches further and highlights gaps in current practices. Ultimately, we aim to facilitate more robust, comparable, and legally aware privacy evaluations in text anonymization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。