发现语言模型中通用的隐私泄露激活方向,可精准放大或抑制个人信息生成。
Discovering Universal Activation Directions for PII Leakage in Language Models
- 通过自生成文本识别模型内部通用的隐私泄露激活方向。
- 沿这些方向调节能显著提升多模型、多数据集的隐私泄露率。
- 无需训练数据即可实现风险放大与缓解,适合安全评估与防护研究。
现代语言模型具有丰富的内部结构,但关于个人身份信息(PII)泄露等隐私敏感行为如何在隐藏状态中表征和调控仍知之甚少。我们提出UniLeak,一种机制解释性框架,识别出模型中的通用激活方向:即在推理时线性添加到残差流中的潜在方向,可一致地提高各类提示下生成PII的概率。这些特定于模型的方向在不同上下文中具有泛化能力,能显著提升PII生成概率,同时对生成质量影响极小。UniLeak无需训练数据或真实PII标注,仅依赖自生成文本即可恢复此类方向。在多个模型和数据集上,沿这些通用方向进行引导,相比现有基于提示的提取方法,显著提升了PII泄露程度。我们的结果揭示了PII泄露的新视角:即模型表示中潜藏信号的叠加,既可放大风险,也可用于风险缓解。
原文摘要 · Abstract (English)
Modern language models exhibit rich internal structure, yet little is known about how privacy-sensitive behaviors, such as personally identifiable information (PII) leakage, are represented and modulated within their hidden states. We present UniLeak, a mechanistic-interpretability framework that identifies universal activation directions: latent directions in a model's residual stream whose linear addition at inference time consistently increases the likelihood of generating PII across prompts. These model-specific directions generalize across contexts and amplify PII generation probability, with minimal impact on generation quality. UniLeak recovers such directions without access to training data or groundtruth PII, relying only on self-generated text. Across multiple models and datasets, steering along these universal directions substantially increases PII leakage compared to existing prompt-based extraction methods. Our results offer a new perspective on PII leakage: the superposition of a latent signal in the model's representations, enabling both risk amplification and mitigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。