大模型在对话中无意泄露用户敏感信息,即使拒绝直接回答也能被还原。
Inadvertent Context Leakage in Language Models

- 通过自适应攻击从普通对话输出中重建上下文秘密
- 2位数字秘密几乎完全复原,4位达82%准确率
- 能力强的模型泄漏更多,适合安全研究者关注
为使AI代理超越简单聊天,需存储用户日程、凭证、健康记录和财务数据等敏感上下文。我们研究这些秘密仅存在于模型上下文窗口时,是否会在模型正常输出中引入隐含关联,导致即便模型正确拒绝直接提取,仍可被重构。进一步研究攻击者能否主动设计提示词,利用模型作为隐蔽载体,通过看似无害的文本传递秘密。两种情况下均采用新型自适应攻击,假设仅有黑盒访问权限。在八个专有模型的受控实验中,发现2位数字秘密可近乎完美重建,4位数字精确匹配率达82%,且均来自模型对常规非对抗性请求的响应。观察到更强模型泄漏更严重:指令遵循能力越强,对上下文秘密越敏感,表明泄漏是能力副产物而非可修复漏洞。我们展示了两种实际攻击:(1) 训练分类器从日常自然语言输出中推断用户记忆的语义谓词(如健康状况、金融事件);(2) 基于强化学习的攻击者从生产级代理中提取完整社会安全号码。
原文摘要 · Abstract (English)
For AI agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere presence of such secrets in a model's context window introduces hidden correlations into the model's benign outputs, allowing reconstruction even when the model correctly refuses direct extraction. We further study whether an adversary can actively engineer prompts that amplify this effect, using the model as a covert carrier to transmit secrets through seemingly innocuous text. In both cases, this limited leakage is exploited using a novel adaptive attack that assumes black-box access to the underlying model. In controlled experiments across eight proprietary models, we find that 2-digit in-context secrets are reconstructed with near-perfect accuracy and 4-digit secrets at 82\% exact match, all from outputs the model produces in response to ordinary, non-adversarial requests. We observe that more capable models leak more: stronger instruction-following amplifies sensitivity to in-context secrets, suggesting leakage is a byproduct of capability as opposed to a patchable bug. We show this leakage enables two practical attacks: (1) a trained classifier that infers semantic predicates about user memories (e.g., health conditions, financial events) from routine natural-language outputs, and (2) an RL-trained adversary that extracts full Social Security Numbers from a production-style agent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。