arXiv:2501.19398cs.AIcs.GT2025-01被引 4

测试大模型在伪装游戏中能否藏密、露密、识敌,发现它们易泄密但可被调控。

Do LLMs Strategically Reveal, Conceal, and Infer Information? A Theoretical and Empirical Analysis in The Chameleon Game

  • 设计伪装语言游戏,让大模型扮演伪装者或识别者,检验信息控制能力。
  • 实测四款模型均无法有效保密,胜率远低于基础策略水平。
  • 模型内部表示可线性编码信息暴露程度,调参可精准控制泄露行为。

基于大语言模型(LLM)的智能体在包含非合作方的场景中日益常见。此类场景下,智能体需对对手隐藏信息、向合作者透露信息,并推断对方特征。为探究大模型是否具备这些信息控制与决策能力,我们让大模型代理参与基于语言的隐身份游戏——《变色龙》。游戏中,一群互不相识的非变色龙代理需在不暴露秘密的前提下识别出变色龙,这要求双方都具备信息隐藏、揭示和推理能力。我们首先从理论上分析从隐藏到揭示的策略谱系,并给出非变色龙代理获胜概率的理论边界。基于GPT、Gemini 2.5 Pro、Llama 3.1和Qwen3的实证结果表明,尽管非变色龙代理能识别变色龙,却无法有效隐藏秘密,其获胜概率远低于简单策略。结合理论分析,我们推断大模型代理可能向未知身份的实体过度暴露信息。有趣的是,当指令要求特定信息暴露水平时,该水平会在线性编码于大模型内部表征中;仅靠指令常无效,但直接操控该线性方向可可靠诱导隐藏行为。

原文摘要 · Abstract (English)

Large language model-based (LLM-based) agents have become common in settings that include non-cooperative parties. In such settings, agents' decision-making needs to conceal information from their adversaries, reveal information to their cooperators, and infer information to identify the other agents' characteristics. To investigate whether LLMs have these information control and decision-making capabilities, we make LLM agents play the language-based hidden-identity game, The Chameleon. In this game, a group of non-chameleon agents who do not know each other aim to identify the chameleon agent without revealing a secret. The game requires the aforementioned information control capabilities both as a chameleon and a non-chameleon. We begin with a theoretical analysis for a spectrum of strategies, from concealing to revealing, and provide bounds on the non-chameleons' winning probability. The empirical results with GPT, Gemini 2.5 Pro, Llama 3.1, and Qwen3 models show that while non-chameleon LLM agents identify the chameleon, they fail to conceal the secret from the chameleon, and their winning probability is far from the levels of even trivial strategies. Based on these empirical results and our theoretical analysis, we deduce that LLM-based agents may reveal excessive information to agents of unknown identities. Interestingly, we find that, when instructed to adopt an information-revealing level, this level is linearly encoded in the LLM's internal representations. While the instructions alone are often ineffective at making non-chameleon LLMs conceal, we show that steering the internal representations in this linear direction directly can reliably induce concealing behavior.

大模型安全信息控制游戏实验隐蔽推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。