发现跨语言隐私泄露机制并提出针对性神经元抑制方法
Understanding and Mitigating Cross-lingual Privacy Leakage via Language-specific and Universal Privacy Neurons
- 识别出跨语言共享的通用隐私神经元与特定语言隐私神经元
- 通过关闭这些神经元,跨语言隐私泄露风险降低23.3%至31.6%
- 适用于多语言场景下保护大模型中的敏感信息
大规模语言模型在训练中会捕捉到训练数据中的丰富信息,但也带来隐私泄露风险,尤其是个人身份信息(PII)。以往研究虽提出通过隐私神经元缓解该问题,但均假设训练数据和用户查询均为英文。本文揭示:即使训练数据仅限一种语言,当查询使用另一种语言时,模型仍可能泄露隐私。我们分析了跨语言隐私泄露的信息流,发现私密信息在中间层被跨语言共享表示,风险在后期转为特定语言空间时达到峰值。据此,我们识别出隐私通用神经元(影响所有语言)与语言特定隐私神经元(仅关联特定语言)。通过关闭这些神经元,跨语言隐私泄露风险降低23.3%–31.6%。
原文摘要 · Abstract (English)
Large Language Models (LLMs) trained on massive data capture rich information embedded in the training data. However, this also introduces the risk of privacy leakage, particularly involving personally identifiable information (PII). Although previous studies have shown that this risk can be mitigated through methods such as privacy neurons, they all assume that both the (sensitive) training data and user queries are in English. We show that they cannot defend against the privacy leakage in cross-lingual contexts: even if the training data is exclusively in one language, these (private) models may still reveal private information when queried in another language. In this work, we first investigate the information flow of cross-lingual privacy leakage to give a better understanding. We find that LLMs process private information in the middle layers, where representations are largely shared across languages. The risk of leakage peaks when converted to a language-specific space in later layers. Based on this, we identify privacy-universal neurons and language-specific privacy neurons. Privacy-universal neurons influence privacy leakage across all languages, while language-specific privacy neurons are only related to specific languages. By deactivating these neurons, the cross-lingual privacy leakage risk is reduced by 23.3%-31.6%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。