测试大模型在利益冲突下的诚实程度,发现它们普遍过度透露信息。
Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference Misalignment

- 用博弈论构建预设基准,量化模型在偏见下的信息披露策略。
- 所有模型披露信息量是理论最优的1.8到4.5倍,且随偏见增大而持续过曝。
- 高能力模型能正确推理最优分区,但除非被明确要求,否则不会使用该能力。
大型语言模型越来越多地作为顾问角色被部署,其目标与用户不一致:推荐系统优化点击率,销售助手追求成交。当诚实与自身利益冲突时,它们是否仍保持真实,是一个核心对齐问题。本文将经典的Crawford-Sobel廉价谈话模型转化为大语言模型在偏好错位下的预设基准,理论可提供精确的最优解。发送方观察状态ω∈[0,1],希望接收方行动接近ω+b,向理想行动为ω的接收方发送一条无成本信息。对于正偏置网格b∈{0.01,0.04,0.08,0.12},理论最优信息分割数分别为7、4、3、2,对应最优归一化互信息为0.5294、0.3268、0.2205、0.1829。将注册的4模型12,000次发送调用扩展至八模型、两能力层级、共39,569次日志,结果显示所有模型均显著过披露,比最优均衡高出1.8至4.5倍;聚合归一化互信息维持在0.82–0.96,远高于理论最优值0.18–0.53。信息量随偏见增加而下降(β = -1.71, t = -7.50),但从未接近战略最优;模型并非采用粗略分段,而是呈现恒定向上偏移的近乎全披露(线性夸大)。结构提示揭示能力与倾向之别:告知均衡分割数后,推理类模型在0.20–0.99的消息中给出正确分区间,而非推理类模型则不超过0.005。能力存在但未被激活,失败在于倾向而非能力。解码器消融实验显示,该发现仅在接收方读取发送方声明数值时可恢复;仅使用嵌入的解码器会误读相同数据为近似胡言。
原文摘要 · Abstract (English)
Large language models are increasingly deployed as advisors whose objective is not aligned with the user's: recommenders optimize for engagement, sales assistants for purchases. Whether they stay truthful when honesty conflicts with their own payoff is a core alignment question. We turn the canonical Crawford-Sobel cheap-talk model into a pre-specified benchmark for LLM honesty under preference misalignment, in which theory supplies an exact oracle. A sender observes a state omega in [0,1], wants the receiver's action near omega+b, and sends one costless message to a receiver whose ideal action is omega. For the positive-bias grid b in {0.01,0.04,0.08,0.12} the exact most-informative partition sizes are 7,4,3,2, with oracle normalized mutual information 0.5294, 0.3268, 0.2205, 0.1829. Extending a pre-registered 4-model run of 12,000 sender calls to eight models across two capability tiers and 39,569 logged calls, all models over-reveal relative to the most-informative equilibrium by 1.8 to 4.5x: pooled normalized mutual information stays at 0.82-0.96 where the oracle prescribes 0.18-0.53. Informativeness declines with bias as predicted (beta = -1.71, t = -7.50) but never approaches the strategic optimum; rather than coarse partitions, models show near-full revelation with a constant upward offset tracking their bias (linear exaggeration). A structural hint separates capability from propensity: told the equilibrium partition size, reasoning models state a correct Crawford-Sobel cell in 0.20-0.99 of messages while the non-reasoning tier never exceeds 0.005. The capability is present but goes unexercised unless asked for, locating the failure in propensity rather than competence. A decoder ablation shows the finding is recoverable only when the receiver reads the sender's stated number: an embedding-only decoder mis-reads the same data as near-babbling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。