训练AI双面间谍,骗对手以为窃取了隐私信息
Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind

- 用强化学习让AI同时理解对手想法并伪装成功
- 结合骗术与心智模型奖励,骗术成功率显著提升
- 适合研究大模型安全与对抗性对话的学者
随着大语言模型成为对话系统的核心,其对对话伙伴意图和状态的推理能力(即心智理论,ToM)对与潜在敌对者安全交互愈发关键。本文提出一项新型隐私导向的ToM挑战——信念引导(ToM-SB),要求防御者作为双面间谍,在共享环境中利用部分先验知识,引导攻击者的信念。为在该任务中取得成功,防御者需建立对攻击者的ToM,并误导其相信已成功获取敏感信息。我们发现,如Gemini3-Pro和GPT-5.4等前沿模型在困难场景下表现不佳,即使经过ToM提示仍难以欺骗攻击者。为此,我们通过强化学习训练模型扮演AI双面间谍,测试仅奖励欺骗成功或仅奖励ToM两种策略。结果发现,欺骗与心智模型能力存在双向涌现关系:单独奖励欺骗可提升ToM,单独奖励ToM亦能增强欺骗能力。在四种不同强度攻击者、六种防御方法及分布内/外评估下,两者性能高度相关,表明信念建模是核心驱动力。结合双重奖励的防御者在硬场景中超越了使用ToM提示的Gemini3-Pro与GPT-5.4。此外,该任务可扩展至更强攻击者,验证了其在分布外设置下的泛化能力与可升级性。
原文摘要 · Abstract (English)
As large language models (LLMs) become the engine behind conversational systems, their ability to reason about the intentions and states of their dialogue partners (i.e., form and use a theory-of-mind, or ToM) becomes increasingly critical for safe interaction with potentially adversarial partners. We propose a novel privacy-themed ToM challenge, ToM for Steering Beliefs (ToM-SB), in which a defender must act as a Double Agent to steer the beliefs of an attacker with partial prior knowledge within a shared universe. To succeed on ToM-SB, the defender must engage with and form a ToM of the attacker, with a goal of fooling the attacker into believing they have succeeded in extracting sensitive information. We find that strong frontier models like Gemini3-Pro and GPT-5.4 struggle on ToM-SB, often failing to fool attackers in hard scenarios with partial attacker prior knowledge, even when prompted to reason about the attacker's beliefs (ToM prompting). To close this gap, we train models on ToM-SB to act as AI Double Agents using reinforcement learning, testing both fooling and ToM rewards. Notably, we find a bidirectionally emergent relationship between ToM and attacker-fooling: rewarding fooling success alone improves ToM, and rewarding ToM alone improves fooling. Across four attackers with different strengths, six defender methods, and both in-distribution and out-of-distribution (OOD) evaluation, we find that gains in ToM and attacker-fooling are well-correlated, highlighting belief modeling as a key driver of success on ToM-SB. AI Double Agents that combine both ToM and fooling rewards yield the strongest fooling and ToM performance, outperforming Gemini3-Pro and GPT-5.4 with ToM prompting on hard scenarios. We also show that ToM-SB and AI Double Agents can be extended to stronger attackers, demonstrating generalization to OOD settings and the upgradability of our task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。