arXiv:2606.11502cs.CLcs.AI2026-06

研究大模型角色扮演时是否真信自己说的话

When Role-playing, Do Models Believe What They Say?

论文配图:When Role-playing, Do Models Believe What They Say?
图 1 · 摘自论文原文
  • 用提示、微调等方法诱导模型扮演不同人物
  • 只有新兴错位训练让模型真正改变信念认知
  • 适用于研究模型内在认知与行为分离的场景

语言模型能说‘地球绕太阳转’,在扮演亚里士多德时却声称相反。现有研究认为人物角色选择是模型行为的核心机制。本研究通过诱导模型扮演与现代共识相悖的角色,采用提示、上下文学习(ICL)、监督微调(SFT)、开放角色训练(OCT)和新兴错位(EM)等多种方法,利用真理探测与行为测试衡量信念内化程度。结果发现信念内化呈现广泛谱系:提示、ICL 和 SFT 仅改变输出,对内部表征影响极小;而 EM 引发显著且广泛的真理表征变化,OCT 在更大模型上表现更明显。理解何时训练改变模型世界观而非仅行为,对日益自主的AI系统至关重要。

原文摘要 · Abstract (English)

Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite. Recent work argues that persona adoption is fundamental to how language models behave, with models selecting the most appropriate persona for a given context. Does such role-playing merely change the model's outputs, or does it also affect what the model internally represents as truthful? We study this question using the role-play of characters whose beliefs differ from the modern consensus, and induce personas with a number of different methods: prompting, in-context learning (ICL), supervised fine-tuning (SFT), and Open Character Training (OCT), and Emergent Misalignment (EM). We measure belief internalization across these approaches with truth probes and with behavioral tests, finding a broad spectrum of belief internalization. Prompting, ICL, and SFT change what the model says with little representational change. EM creates a large, broad shift in the model's truth representation, and OCT a smaller shift that is clearest on the larger model. Understanding when training changes a model's worldview rather than merely its behavior may become increasingly important as AI systems are entrusted with greater autonomy and influence.

角色扮演信念内化大模型认知行为一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。