梳理文本匿名化最新进展,兼顾隐私保护与数据可用性。
A Survey on Current Trends and Recent Advances in Text Anonymization
- 以命名实体识别为基础,结合大模型实现智能匿名
- 提出多领域定制方案,应对医疗、金融等场景挑战
- 强调隐私-效用平衡,适合研究者与从业者参考
文本数据中敏感个人信息的泛滥要求强有力的匿名化技术以保护隐私并满足法规要求,同时保持数据在各类下游任务中的可用性。本综述全面概述了当前文本匿名化技术的趋势与最新进展。首先讨论以命名实体识别为核心的基线方法,随后分析大语言模型带来的变革,揭示其作为高效匿名工具及强大去匿名化威胁的双重角色。进一步探讨医疗、法律、金融和教育等关键领域的特定挑战与定制化解决方案。研究涵盖形式化隐私模型与风险感知框架的先进方法,以及作者身份匿名化的专门子领域。此外,还回顾了评估框架、综合度量标准、基准测试集与实用工具包,支持真实世界部署。本文整合现有知识,识别新兴趋势与持续挑战,包括不断演变的隐私-效用权衡、准标识符处理需求以及大模型能力的影响,旨在为学术界与实践界未来研究提供指引。
原文摘要 · Abstract (English)
The proliferation of textual data containing sensitive personal information across various domains requires robust anonymization techniques to protect privacy and comply with regulations, while preserving data usability for diverse and crucial downstream tasks. This survey provides a comprehensive overview of current trends and recent advances in text anonymization techniques. We begin by discussing foundational approaches, primarily centered on Named Entity Recognition, before examining the transformative impact of Large Language Models, detailing their dual role as sophisticated anonymizers and potent de-anonymization threats. The survey further explores domain-specific challenges and tailored solutions in critical sectors such as healthcare, law, finance, and education. We investigate advanced methodologies incorporating formal privacy models and risk-aware frameworks, and address the specialized subfield of authorship anonymization. Additionally, we review evaluation frameworks, comprehensive metrics, benchmarks, and practical toolkits for real-world deployment of anonymization solutions. This review consolidates current knowledge, identifies emerging trends and persistent challenges, including the evolving privacy-utility trade-off, the need to address quasi-identifiers, and the implications of LLM capabilities, and aims to guide future research directions for both academics and practitioners in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。