提出新指标评估语言模型中上下文隐私泄露风险
Estimating Privacy Leakage of Augmented Contextual Knowledge in Language Models
- 基于差分隐私设计上下文影响度量,分离模型参数知识
- 发现当上下文偏离模型预训练分布时隐私泄露明显
- 适用于评估大模型问答中的隐私风险,指导安全应用
语言模型在问答等任务中依赖参数化知识与上下文知识的结合,但上下文可能包含隐私信息,其泄露风险尚不明确。直接对比输出与上下文会高估风险,因模型本身可能已掌握该知识。为此,我们提出‘上下文影响’(context influence)这一基于差分隐私的度量,有效分离模型参数知识,量化各上下文子集对生成结果的影响。实验表明,当上下文与模型参数知识分布不一致时,隐私泄露发生。我们验证了该指标能准确归因于增强上下文,并评估了模型规模、上下文大小、生成位置等因素对隐私泄露的影响。结果可为实际应用中上下文增强的安全性提供指导。
原文摘要 · Abstract (English)
Language models (LMs) rely on their parametric knowledge augmented with relevant contextual knowledge for certain tasks, such as question answering. However, the contextual knowledge can contain private information that may be leaked when answering queries, and estimating this privacy leakage is not well understood. A straightforward approach of directly comparing an LM's output to the contexts can overestimate the privacy risk, since the LM's parametric knowledge might already contain the augmented contextual knowledge. To this end, we introduce *context influence*, a metric that builds on differential privacy, a widely-adopted privacy notion, to estimate the privacy leakage of contextual knowledge during decoding. Our approach effectively measures how each subset of the context influences an LM's response while separating the specific parametric knowledge of the LM. Using our context influence metric, we demonstrate that context privacy leakage occurs when contextual knowledge is out of distribution with respect to parametric knowledge. Moreover, we experimentally demonstrate how context influence properly attributes the privacy leakage to augmented contexts, and we evaluate how factors -- such as model size, context size, generation position, etc. -- affect context privacy leakage. The practical implications of our results will inform practitioners of the privacy risk associated with augmented contextual knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。