arXiv:2602.07164cs.CLcs.AI2026-02被引 2

发现大模型参数中隐藏着性格子网络,无需训练即可激活不同人格。

Your Language Model Secretly Contains Personality Subnetworks

  • 通过小数据集识别不同人格的激活模式,用掩码分离出轻量级子网络。
  • 在多种评测中,子网络人格对齐度显著优于依赖外部知识的方法。
  • 提出对比剪枝法,可区分对立人格(如内向-外向),适合个性化控制研究者。

人类会根据社交情境切换不同人格。大型语言模型(LLMs)也表现出类似的行为灵活性。现有方法通常通过提示、检索增强生成(RAG)或微调等外部知识实现行为适应。我们提出:是否真的需要外部上下文或参数来调整行为,还是这些知识早已嵌入模型参数中?本工作表明,LLMs 的参数空间中已存在专门化的人格子网络。利用少量校准数据集,我们识别出与不同人格相关的独特激活特征。基于这些统计特征,我们设计了一种掩码策略,以隔离轻量级人格子网络。在此基础上,进一步探讨如何从模型中发现导致二元对立人格(如内向-外向)的相反子网络。为增强二元对立场景下的分离效果,我们引入对比剪枝策略,识别出造成对立人格统计差异的参数。该方法完全无需训练,仅依赖模型现有的参数空间。在多种评估设置下,所得子网络的人格对齐度显著优于需外部知识的基线方法,且更高效。研究结果表明,多样化的类人行为并非仅由外部诱导,而是已嵌入于模型参数空间中,为可控、可解释的个性化提供了新视角。

原文摘要 · Abstract (English)

Humans shift between different personas depending on social context. Large Language Models (LLMs) demonstrate a similar flexibility in adopting different personas and behaviors. Existing approaches, however, typically adapt such behavior through external knowledge such as prompting, retrieval-augmented generation (RAG), or fine-tuning. We ask: do LLMs really need external context or parameters to adapt to different behaviors, or do they already have such knowledge embedded in their parameters? In this work, we show that LLMs already contain persona-specialized subnetworks in their parameter space. Using small calibration datasets, we identify distinct activation signatures associated with different personas. Guided by these statistics, we develop a masking strategy that isolates lightweight persona subnetworks. Building on the findings, we further discuss: how can we discover opposing subnetwork from the model that lead to binary-opposing personas, such as introvert-extrovert? To further enhance separation in binary opposition scenarios, we introduce a contrastive pruning strategy that identifies parameters responsible for the statistical divergence between opposing personas. Our method is entirely training-free and relies solely on the language model's existing parameter space. Across diverse evaluation settings, the resulting subnetworks exhibit significantly stronger persona alignment than baselines that require external knowledge while being more efficient. Our findings suggest that diverse human-like behaviors are not merely induced in LLMs, but are already embedded in their parameter space, pointing toward a new perspective on controllable and interpretable personalization in large language models.

人格建模子网络无训练可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。