arXiv:2607.26853cs.CLcs.AI2026-07

发现大模型内部存在可操控的人格特征表示,能稳定影响行为表现。

From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs

论文配图:From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs
图 1 · 摘自论文原文
  • 通过对比行为对分析内部特征,识别出人格极性对应的稀疏表示。
  • 在不同情境下干预特征,可双向调控行为表现且保持回应合理性。
  • 在社交任务中验证行为变化符合人类人格研究规律,具有实证意义。

人类人格理论认为特质并非单一评分,而是个体在人、情境与行为互动中的稳定倾向。现有大模型人格研究多聚焦人格条件下的输出,缺乏对内部人格表征、跨情境表达及其如何塑造行为的机制证据。基于Funder人格三元框架,我们构建分析体系:将‘人’定义为人格相关内部表示,‘情境’为诱发特质响应的上下文,‘行为’为更广泛社交任务中的反应模式。首先,利用共享情境下的对比行为对,通过SAE分解识别出对应人格两极的稀疏内部特征,并通过行为-情境效应、词级激活模式和抗改写鲁棒性验证其人格相关性。其次,特征级干预可在独立多样的情境中引发双向人格化行为转变,同时保持回应有效性,证明跨情境一致性。最后,将相同干预应用于社交智能任务,观察到行为变化呈现与人类研究一致的收益-代价权衡模式,提供行为层面的验证。结果表明,大模型包含可操控的类人格表征,连接内部状态、情境表达与行为结果。

原文摘要 · Abstract (English)

Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, situations, and behaviors. Existing studies of personality-related behavior in LLMs have primarily focused on outputs elicited under personality conditioning, characterizing observable trait-related expressions while lacking mechanistic evidence for the existence of internal personality-related representations, their cross-situational expression, and how these representations shape specific behaviors. Building on Funder's personality triad framework, we adapt its three components for LLM analysis: Person as personality-related internal representations, Situation as contexts that afford trait-relevant responses, and Behavior as response patterns on broader social tasks. We introduce a framework for discovering, controlling, and validating trait-like representations in LLMs. First, using contrastive behavior pairs grounded in shared situations, we identify sparse internal features associated with opposing poles of personality traits through SAE decomposition. We validate their trait relevance through effects on behavior to situation, token-level activation patterns, and robustness to paraphrasing. Second, feature-level interventions induce bidirectional trait-related shifts across a separate, diverse set of situations while preserving response validity, demonstrating consistent expression across contexts. Third, applying the same interventions to social intelligence tasks reveals behavioral changes with benefit-tradeoff patterns consistent with findings from human personality research, providing behavioral-level validation beyond personality scores. Our findings provide evidence that LLMs contain controllable trait-like representations linking internal states, situational expression, and behavioral outcomes.

大模型人格表征分析行为建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。