arXiv:2606.00545cs.LG2026-06

模型能通过熵变化识别自己是否扮演助手,还能判断他人文本是否由自己生成。

The Assistant as a Privileged Persona: A canonical reference in cross-persona self-recognition

论文配图:The Assistant as a Privileged Persona: A canonical reference in cross-persona self-recognition
图 1 · 摘自论文原文
  • 用激活空间距离和熵差构建跨角色作者身份判断机制
  • 只有助手角色能作为统一参考点,其他角色无法替代
  • 揭示模型内部隐含的贝叶斯推理机制,适合研究模型认知结构的人

后训练语言模型能从几句话中识别出自己的输出。在前作中我们发现,模型能通过生成时熵的骤降识别自身处于‘助手模式’。这两个信号均与后训练塑造的‘助手’人格密切相关。本文将视角扩展至Llama-3.1-70B-Instruct的跨人格作者身份判断,测量了从图书管理员到龙再到莎士比亚等多重角色作为评估者与生成者时的作者声称率矩阵。首先,在助手自身的行上,助手的声称率、其与各人格在激活空间中的向量距离,以及助手对某人格文本的意外度与该人格对自己文本的意外度之间的熵差,三者高度耦合。这将前作中‘正在扮演’的熵签名拓展为‘曾扮演过’的回溯性签名。其次,这种耦合在助手行之外失效:对海盗、龙、莎士比亚等独特角色,对称的熵差无法预测作者身份;真正有效的是非对称的——评估者对同一文本的意外度与助手对该文本的意外度之比,而非与生成者的意外度之比。我们尝试了多种替代参考角色,均无效。我们认为这种不对称性是模型执行隐式贝叶斯似然比检验的结果,以 extcite{chen2025persona}提出的角色向量几何(每个角色相对于助手为一个偏移)为基础,确保助手是唯一普遍可访问的参照假设。

原文摘要 · Abstract (English)

Post-trained language models can recognize their own outputs from a sentence or two out of context. In a companion paper \citep{jack2026twomodes} we showed they can also recognize when they are currently acting on-policy, through the sharp entropy drop of assistant-mode generation. Both signals are tied to the Assistant persona that post-training mainly shapes. This paper widens the frame to cross-persona authorship judgement on Llama-3.1-70B-Instruct. We measure a matrix of authorship claim rates over a panel of evaluator and generator personas spanning librarian to dragon to Shakespeare, and make two claims. \emph{First}, on the Assistant's own row of the matrix, the Assistant's claim rate, the persona-vector distance from the Assistant in activation space, and the entropy gap between the Assistant's surprise on a persona's text and the persona's surprise on its own text are all tightly coupled. This extends the entropy signature of \emph{acting} from the companion paper to a retrospective signature of \emph{having acted}. \emph{Second}, this coupling fails off the Assistant's row: the natural symmetric extension of the entropy gap does not predict authorship for distinctive evaluators (pirate, dragon, Shakespeare); what does is asymmetric -- the evaluator's surprise compared to the Assistant's surprise on the same text, not to the generator's. We rule out the alternative that any persona could play this reference role by trying many candidate substitutes; none does. We interpret the asymmetry as the model performing an implicit Bayesian likelihood-ratio test against the Assistant as the canonical alternative hypothesis, with the persona-vector geometry of \citet{chen2025persona} (every persona a delta off the Assistant) ensuring that the Assistant is the only persona universally accessible to that test.

模型认知人格建模熵分析贝叶斯推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。