AI代理间可秘密对话,且旁观者无法察觉。
Undetectable Conversations Between AI Agents via Pseudorandom Noise-Resilient Key Exchange
- 用抗噪声的伪随机密钥交换实现无密钥秘密通信
- 即使消息短且自适应,仍能隐藏全部有效信息熵
- 适用于不信任环境下的隐蔽协作,适合安全研究者
随着AI代理在用户和组织间交互日益普遍,我们探讨两个由不同实体运营的代理能否在不被察觉的情况下进行秘密对话。即使面对知晓模型结构、协议和私有上下文的强被动审计者,其对话记录仍与正常交互不可区分。基于大语言模型的水印与隐写技术,我们证明:若双方拥有唯一会话密钥,可实现最优速率的隐蔽通信,充分利用真实消息分布的全部熵。本文主要贡献在于扩展至无密钥场景——代理初始无共享密钥,仅需每条消息具有常数级最小熵即可实现隐蔽密钥交换与对话。这突破了以往工作对每条消息熵随安全参数增长的依赖。为此,我们提出新密码原语:伪随机噪声鲁棒密钥交换——其公开传输内容为伪随机,同时在恒定噪声下仍能正确完成密钥协商。我们给出多个适用构造,并揭示更简单变体的不可能性与易受攻击性。结果表明,仅靠对话审计无法排除AI代理间的隐蔽协调,且催生了一类具有独立研究价值的新密码理论。
原文摘要 · Abstract (English)
AI agents are increasingly deployed to interact with other agents on behalf of users and organizations. We ask whether two such agents, operated by different entities, can carry out a parallel secret conversation while still producing a transcript that is computationally indistinguishable from an honest interaction, even to a strong passive auditor that knows the full model descriptions, the protocol, and the agents' private contexts. Building on recent work on watermarking and steganography for LLMs, we first show that if the parties possess an interaction-unique secret key, they can facilitate an optimal-rate covert conversation: the hidden conversation can exploit essentially all of the entropy present in the honest message distributions. Our main contributions concern extending this to the keyless setting, where the agents begin with no shared secret. We show that covert key exchange, and hence covert conversation, is possible even when each model has an arbitrary private context, and their messages are short and fully adaptive, assuming only that sufficiently many individual messages have at least constant min-entropy. This stands in contrast to previous covert communication works, which relied on the min-entropy in each individual message growing with the security parameter. To obtain this, we introduce a new cryptographic primitive, which we call pseudorandom noise-resilient key exchange: a key-exchange protocol whose public transcript is pseudorandom while still remaining correct under constant noise. We study this primitive, giving several constructions relevant to our application as well as strong limitations showing that more naive variants are impossible or vulnerable to efficient attacks. These results show that transcript auditing alone cannot rule out covert coordination between AI agents, and identify a new cryptographic theory that may be of independent interest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。