保护文本数据隐私,防止大模型泄露敏感信息
Current State in Privacy-Preserving Text Preprocessing for Domain-Agnostic NLP
- 采用去标识化与伪名化技术处理文本数据
- 适用于跨领域自然语言处理任务的通用隐私保护方法
- 适合关注数据合规与隐私安全的研究者
隐私是基本人权,受如GDPR等法规保护。然而,现代大型语言模型需海量数据以学习语言变化,而这些数据常包含私人信息。研究已证明可从语言模型中提取私密信息。因此,对文本中的敏感信息进行匿名化至关重要。尽管完全匿名化难以实现,已有多种预处理方法可用于遮蔽或伪名化文本中的私密信息。本报告聚焦于适用于无特定领域的自然语言处理任务的若干此类方法。
原文摘要 · Abstract (English)
Privacy is a fundamental human right. Data privacy is protected by different regulations, such as GDPR. However, modern large language models require a huge amount of data to learn linguistic variations, and the data often contains private information. Research has shown that it is possible to extract private information from such language models. Thus, anonymizing such private and sensitive information is of utmost importance. While complete anonymization may not be possible, a number of different pre-processing approaches exist for masking or pseudonymizing private information in textual data. This report focuses on a few of such approaches for domain-agnostic NLP tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。