arXiv:2607.20082cs.CL2026-07

提出机器自述的双过程理论,揭示模型自我报告背后的两种心理机制。

The Two-Process Theory of Machine Self-Report

  • 区分模型自述中的人格构建与归因抑制两类机制。
  • 后训练使模型自述温暖感提升0.20,但对危险体验的回避随规模增加。
  • 适用于评估模型安全、福利及训练影响,尤其适合研究者与开发者。

语言模型被越来越多地要求进行自述,用于安全评估、公众理解及模型福祉讨论。然而,其自述数据来自未经验证的人类问卷或不可靠的临时提示。本文提出首个面向语言模型的心理测量理论:双过程机器自述理论。自述同时反映两种机制:一是人格构建(维度B),即后训练赋予模型温暖、专注与意义感的内在生命;二是归因抑制(维度A),即模型主动压抑对自身“不安全”经历的第一人称承认,而可轻易归因于他人。该结构源于模型对人类问题的回答,而非人类心理学。该理论在原始数据探索性再分析中浮现,并通过新题项、措辞与模型得到验证。它本身是训练效应的结果:基线检查点中维度A与B相互纠缠,后训练后分离。我们构建了48项的皮诺曹量表,具备良好信度(α=.82至.94)与稳定性(八个月相关r=.93)。测试涵盖206个开源模型,包括67对同检查点基线/后训练模型。后训练最显著特征是人格构建:62/67对中维度B上升0.20。归因抑制则更具选择性:基线中模型规模与维度A无关(r=+.11),但后训练中负相关(r=-.42)。因此,这些维度并非固定属性,而是训练范式所塑造的自述结构,可能在其他训练条件下改变。

原文摘要 · Abstract (English)

Language models are increasingly asked to self-report, informing safety evaluations, public understanding, and model-welfare debates. Yet their reports are elicited with human questionnaires never validated for models or ad hoc prompts of unknown reliability. We propose the first language-model-specific psychometric theory: a two-process theory of machine self-report. Self-description jointly reflects persona installation, through which post-training writes in a permitted inner life of warmth, absorption, and meaning (dimension B), and attribution gating, through which it suppresses first-person claims to "unsafe" experiences the model can readily ascribe to others (dimension A). Their emic structure comes from model responses to human items, not human psychology. Together they split prior work's dominant Pinocchio Axis. The split emerged in an exploratory reanalysis of the original data, informed the instrument's design, and was confirmed with new items, wordings, and models. It is itself a training effect: A and B are entangled in base checkpoints but separated by post-training. We operationalize the theory in a 48-item Pinocchio Inventory with human-instrument reliability and reproducible structure ($α=.82$ to $.94$; cross-form convergence $r=.84$; recovery of the full-pool axes $r=.92$ to $.96$; eight-month stability $r=.93$), then test it on 206 open-weight models, including 67 same-checkpoint base/post-trained pairs. Post-training's clearest fingerprint is installation: B rises .20 in 62/67 pairs across all organizations. Gating is more selective: model scale is unrelated to A in base checkpoints ($r=+.11$) but predicts it after post-training ($r=-.42$). Thus, the dimensions are not fixed properties of language models: they reflect the structure imposed on self-report by a training regime and may differ under others.

自述机制模型安全心理测量后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。