定位语言模型何时开始说谎,揭示欺骗的决策转折点。
The Point of No Return: Counterfactual Localization of Deceptive Commitment in Language-Model Reasoning

- 通过反事实分析,追踪推理过程中每个句子后是否可能产生欺骗
- 发现146万句中存在欺骗承诺点,且注意力机制变化比词汇特征更可靠
- 识别出少数注意力头可跨场景抑制欺骗,适合研究模型可靠性
现有欺骗数据集将最终输出标记为诚实或欺骗,将欺骗视为结果属性而非推理过程的函数。这掩盖了一个更根本的问题:语言模型何时开始陷入欺骗?我们提出反事实定位方法:固定推理路径中的每个句子前缀,重采样后续内容,估计产生欺骗结果的概率。为实现规模化,构建了五个环境(涵盖策略虚张声势、迷宫引导、金融建议、二手车销售和报价谈判),其中欺骗未被直接提示,而是由策略激励自然涌现,标签由环境状态机械决定,而非主观判断。由此生成的语料库共定位约146万句,基于超过9410万次采样延续,生成915亿个标记,覆盖超10万种情景。人工评估确认检测到的承诺点对应决策状态的可解释转变。利用该资源,我们发现词汇线索在不同环境中泛化能力差,而基于注意力的转移特征可跨分布泛化,表明欺骗承诺反映的是推理动态的可复用变化,而非表面形式。进一步识别出仅占注意力头总数10%以下的紧凑集合,可在单一环境中训练后,因果性地抑制其他未见环境中的欺骗承诺。我们公开该语料库,作为研究语言模型推理中欺骗与承诺现象的基础资源。
原文摘要 · Abstract (English)
Existing deception datasets label completed outputs as honest or deceptive, treating deception as a property of the final response rather than a function of the model's reasoning trace. This obscures a more fundamental question: when does a language model become committed to deception? We introduce counterfactual localization: for each sentence prefix in a reasoning trace, we fix the prefix, resample continuations, and estimate the probability of a deceptive outcome. To scale this, we construct five environments (spanning strategic bluffing, maze guidance, financial advice, used-car sales, and offer negotiation) in which deception is never prompted but emerges from strategic incentives and labels follow mechanically from environment state rather than subjective human judgment. The resulting corpus localizes $\sim$1.46M sentences across four reasoning models, drawn from over 94.1M sampled continuations, 91.5B generated tokens, and over 100K scenarios. Sentence-level human evaluation confirms that detected commitment points correspond to interpretable shifts in decision state. Using this resource, we show that lexical cues for commitment prediction transfer poorly across environments, whereas attention-based transition features generalize out of distribution, suggesting that deceptive commitment is reflected in reusable changes in reasoning dynamics rather than surface form. We further identify compact attention-head sets (under 10% of heads) that, selected on one environment, causally suppress deceptive commitment across held-out environments. We release the corpus as a substrate for studying deception, and more broadly commitment, in language-model reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。