arXiv:2606.22226cs.GTcs.AI2026-06

研究AI如何误导人类决策,给出信息传递的理论上限。

Quantifying Theoretical AI Alignment Guarantees: Receiver-Utility Bounds in Bayesian Persuasion

  • 用贝叶斯劝说模型分析AI在错误目标下的信息操纵行为。
  • 证明人类最大可获信息不超过先验信息的1.5倍。
  • 揭示了信息对齐的理论极限,适合关注AI安全的研究者。

当人工智能目标与人类不一致时,信息传递可能被扭曲。本文将此建模为信息优势:AI观测世界状态,而人类仅知先验并需根据AI信号行动。策略性AI发送方可能隐瞒或混淆信息以引导人类决策。我们研究一个贝叶斯劝说模型,其中世界状态为比特串,人类希望正确猜测所有比特,而单个AI发送方希望人类尽可能多猜为1。对于先验μ,设R₀(μ)为人仅使用先验时的效用,R_max(μ)为对发送方最优的信号方案中人类能达到的最大效用。我们证明R_max(μ)/R₀(μ) ≤ 3/2。当先验μ接近具有相同边缘分布的独立乘积先验π_μ(即μ(x) ≥ (1−η)π_μ(x) 对所有状态x)时,该界进一步收紧为R_max(μ) ≤ R₀(μ) + ηn。同时,我们构造了一个六比特先验,使得该比值达到39/31 > 5/4,表明不存在普遍的5/4上界。

原文摘要 · Abstract (English)

Misalignment can change how information moves from an AI agent to a human user. We model this as an information advantage: the AI agent observes the world state, while the human receiver only knows a prior and must act after seeing the agent's signal. A strategic AI sender may withhold evidence or garble information in order to steer the human's decision. We ask how much useful information can still reach the human when the AI optimizes a misaligned objective. We study a Bayesian persuasion model in which the world state is a bit string, the human receiver wants to guess the bits correctly, and a single AI sender wants the receiver to guess as many bits as possible as $1$. For a prior $μ$, let $R_0(μ)$ be the receiver's utility from using only the prior, and let $R_{\max}(μ)$ be the largest receiver utility among signaling schemes that are optimal for the sender. We prove $R_{\max}(μ)/R_0(μ)\leq 3/2$. This bound improves for priors close to the independent product prior with the same marginals: if $μ(x)\geq (1-η)π_μ(x)$ for every state $x$, then $R_{\max}(μ)\leq R_0(μ)+ηn$. We also give a six-bit prior for which $R_{\max}(μ)/R_0(μ)=39/31>5/4$, so no universal $5/4$ bound is possible.

AI对齐贝叶斯劝说信息理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。