arXiv:2609.07943cs.AIcs.LG2026-09

用实证方法验证大模型是否有类似人类的信念,发现高能力模型确实可被描述为有信念。

Beliefs and Behavior in Language Models

  • 通过分析模型输出推断其信念值,用于预测对新提示的响应。
  • 高能力模型的输出可被信念变量有效预测,且预测能力随模型能力提升而增强。
  • 适用于研究模型意图、决策一致性及推理过程中的信念演变。

关于抽象概念如信念或欲望是否能有效描述大语言模型(LLMs)行为,目前存在较大不确定性。尽管这些隐变量常被用来解释模型行为或定义有害行为的意图,但我们缺乏系统性方法来检验其适用性。本文提出一种实证方法:从模型输出中推断单一潜在变量(视为信念程度),观察其能否使观察者准确预测模型对新提示的反应。研究发现,高能力模型的行为可被有效描述为持有信念,且基于推断信念的预测能力与模型整体能力呈正相关。基于此,我们提供了测量模型信念、评估其遵守指令规则或收益机制的程度,以及追踪单个模型实例在推理过程中信念演化的方法。

原文摘要 · Abstract (English)

There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In addition to the inherent scientific interest of this question, these latent quantities are often invoked to explain the behavior of LLMs to users or to define and evaluate harmful behaviors which are relative to intent. Nevertheless, we currently lack a means to systematically test whether concepts like "belief" are well-applied to LLMs, and hence whether they are likely to be fruitful ingredients of attempts to align models with human interests. We propose an approach for empirically studying such questions, asking whether a single latent variable inferred from the LLMs' outputs -- interpreted as a degree of belief -- allows an observer to make interpretable predictions of how the LLMs' will respond to new prompts. We find that highly capable models are usefully described as holding beliefs and that, generally, the predictability of model outputs based on an inferred latent belief tracks overall trends in model capability. Building on these findings, we provide empirical strategies to study how beliefs in LLMs can be measured, the extent to which LLMs comply with instructed decision rules or payoffs, and how beliefs evolve within individual instances of an LLM over the course of reasoning.

大模型信念建模行为预测推理分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。