通过对比激活定位特定人格神经元,实现精准可控的人格编辑。
DPN-LE: Dual Personality Neuron Localization and Editing for Large Language Models

- 对比高/低人格样本的MLP激活,定位专属人格神经元。
- 仅干预0.5%神经元,保持推理能力且人格控制效果优秀。
- 适合需要精细调整模型人格、避免性能下降的研究者。
随着大语言模型的广泛应用,理解其人格表征机制变得至关重要。现有个性化编辑方法多依赖神经元编辑,需修改大量神经元,导致性能显著下降。本文探究并量化了这一问题:1)现有方法虽能改变人格,但会降低整体能力;2)神经元具有多功能性,关联人格特质与通用知识;3)对立人格特征呈现明显互斥的表征模式。基于此,提出DPN-LE(双人格神经元定位与编辑),通过对比高人格与低人格样本的MLP激活,构建逐层控制向量,并结合Cohen's $d$效应量与激活幅度进行双准则筛选,识别出互斥神经元子集。对这些稀疏神经元进行线性干预,可在推理时实现精确人格调控。仅用每种人格1,000对对比样本,干预约0.5%神经元,即在LLaMA-3-8B-Instruct和Qwen2.5-7B-Instruct上实现良好人格控制与显著更优的能力保留。实验验证了该方法的有效性与泛化性。
原文摘要 · Abstract (English)
With the widespread adoption of large language models (LLMs), understanding their personality representation mechanisms has become critical. As a novel paradigm in Personality Editing, most existing methods employ neuron-editing to locate and modify LLM neurons, requiring changes to numerous neurons and leading to significant performance degradation. This raises a fundamental question: Are all modified neurons directly related to personality representation? In this work, we investigate and quantify this specificity through assessments of general capability impact and representation-level patterns. We find that: 1) Current methods can change personalities but reduce overall performance. 2) Neurons are multifunctional, connecting personality traits and general knowledge. 3) Opposing personality traits demonstrate distinctly mutually exclusive representation patterns. Motivated by these findings, we propose DPN-LE (Dual Personality Neuron Localization and Editing), which identifies personality-specific neurons by contrasting MLP activations between high-trait and low-trait samples. DPN-LE constructs layer-wise steering vectors and applies dual-criterion filtering based on Cohen's $d$ effect size and activation magnitude to isolate mutually exclusive neuron subsets. Sparse linear intervention on these neurons enables precise personality control at inference time. Using only 1,000 contrastive sample pairs per trait, DPN-LE intervenes on $\sim$0.5\% of neurons while achieving competitive personality control and substantially better capability preservation across reasoning tasks. Experiments on LLaMA-3-8B-Instruct and Qwen2.5-7B-Instruct demonstrate the effectiveness and generalizability of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。