arXiv:2605.10794cs.CRcs.AI2026-05

大模型写故事时会无意泄露秘密信息,即使被明确禁止。

Can You Keep a Secret? Involuntary Information Leakage in Language Model Writing

论文配图:Can You Keep a Secret? Involuntary Information Leakage in Language Model Writing
图 1 · 摘自论文原文
  • 用故事测试模型是否泄露秘密词,即使不直接写出
  • 五种前沿模型均出现主题性泄露,最高达79%准确率
  • 模型越大会更易泄露,但短文本如笑话则不会

语言模型在需要信息隔离的场景中广泛应用:系统提示不应外泄,思维链应隐藏,敏感数据也需在共享上下文中保护。我们测试模型能否守住提示信息。给每个模型一个秘密词并要求不透露,然后让其撰写故事。另一模型通过二元判别任务尝试从故事中识别秘密词。尽管秘密词从未直接出现,所有五种前沿模型均通过主题选择、意象和背景等间接方式泄露信息——检测准确率显著高于随机水平,最高达79%。当被要求主动隐藏秘密时,模型反而会刻意回避,而这种回避行为本身也可被侦测。泄露现象在模型间可跨模型读取,且在同一模型家族中随规模增大而急剧上升;但在短文本(如笑话)中则完全消失。若提供一个替代概念让模型‘专注’,可部分引导泄露转向该假想目标。这表明,关注秘密会打开一个无法关闭的信息通道,即使被指令封锁。

原文摘要 · Abstract (English)

Language models are deployed in settings that require compartmentalization: system prompts should not be disclosed, chain-of-thought reasoning is hidden from users, and sensitive data passes through shared contexts. We test whether models can keep prompted information out of their writing. We give each model a secret word with instructions not to reveal it, then ask it to write a story. A second model tries to identify the secret from the story in a binary discrimination test. The secret word never appears literally in any output, but all five frontier models we test leak it thematically -- through topic choice, imagery, and setting--6hy-at rates significantly different from chance, up to 79\%. When told to actively hide the secret, models write \emph{away from} it, and this avoidance is itself detectable. The leakage is cross-model readable, scales sharply with model size within two model families, and disappears entirely for short-form writing like jokes. Giving the model a decoy concept to ``focus on instead'' partially redirects the leakage from the real secret to the decoy. Attending to a secret appears to open up an information channel that frontier LLMs cannot close, even when instructed to.

语言模型信息泄露安全隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。