arXiv:2605.25151cs.AIcs.CE2026-05被引 1

测试大模型是否真有类似人类的风险决策机制。

Representation Without Control: Testing the Realization Effect in Language Models

  • 用行为经济学中的实现效应检验模型决策逻辑。
  • 模型内部存在可解码的实现状态信号,但无法操控风险选择。
  • 模型行为敏感、可读取、能控制三者不必然共现,需谨慎解读内部信号。

大型语言模型被越来越多地用作行为模拟器,但其输出何时反映类人认知机制,而非仅受提示影响的表面模式仍不明确。本文通过实现效应——行为经济学中一个经典发现,即纸面盈亏与实际盈亏后的风险偏好差异——来研究这一问题。我们从三个层面评估大模型行为:仅依赖提示的行为敏感性、内部表示的线性解码能力,以及通过激活操控实现因果控制。结果显示,模型对提示具有系统性敏感性,但方向不符合人类实现效应预测。Gemma 模型在第 18 层的残差流中存在可线性解码的实现状态信号,且在未见提示上具有泛化能力。然而,沿该方向操控激活,并未可靠改变下游风险决策,此结果在正负尺度及符号对称实验中均成立。行为敏感性、潜在信号解码和因果控制是三个独立属性,成功解码不代表模型在决策中真正依赖该表征。

原文摘要 · Abstract (English)

Large language models are increasingly used as behavioral simulators, but it remains unclear when their outputs reflect human-like cognitive mechanisms rather than prompt-sensitive surface patterns. We study this question through the realization effect, a well-characterized finding in behavioral economics in which risk-taking differs systematically after paper versus realized gains and losses. We evaluate LLM behavior at three levels: prompt-only behavioral sensitivity, linear readout of internal representations, and causal control via activation steering. Prompt-only results show systematic condition sensitivity, but the directional pattern does not reproduce human realization-effect predictions. Gemma's residual stream contains a linearly decodable realization-status signal at layer 18 that generalizes to held-out prompts. Steering along this direction does not, however, reliably shift downstream risk choices, a null result that holds across positive scales and in a negative sign-symmetry run. Behavioral sensitivity, latent readout, and causal control are three distinct properties that do not automatically co-occur, and successful latent readout is insufficient evidence that a model behaviorally relies on a representation during downstream decision-making.

大模型行为认知机制实现效应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。