arXiv:2606.28770cs.AI2026-06

通过干预模型隐空间特征,精准控制大模型人格表现

Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions

论文配图:Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions
图 1 · 摘自论文原文
  • 用稀疏自编码器定位人格相关隐层特征方向
  • 微小特征偏移即可增强目标人格特质且不损语言能力
  • 适合想精细调控AI性格的研究者与开发者

大型语言模型(LLMs)已能生成体现人类OCEAN人格特质的文本。以往方法多依赖提示工程或微调来塑造人格。本文提出一种机制可解释性方法,直接干预模型的隐层特征。通过稀疏自编码器(SAEs)和对比激活分析,识别残差流中对应目标OCEAN特质的隐层方向,并在激活空间中构建加性引导向量。实验表明,对隐藏状态施加微小加性偏移,可有效增强目标人格特质,同时保持原有语言建模性能。为确定最优特征偏移组合,采用线性加权启发式与网格搜索优化,在人格表达与任务性能间取得平衡。该方法在保持高基准性能的前提下,实现了人格特质的可控机制级调节。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated the ability to simulate human-like OCEAN personality traits in generated text. Previous efforts have focused on prompt engineering or fine-tuning to shape LLM personality. In this work, we propose a mechanistic interpretability approach that directly intervenes on the model's latent features. Our method identifies latent directions in the residual stream corresponding to a target OCEAN trait using sparse autoencoders (SAEs) and contrastive activation analysis. We formalize an additive steering vector in activation space and demonstrate how applying a small additive shift to the hidden states enhances the target trait while preserving overall language modeling performance. To determine the optimal combination of feature shifts, we explore a linear weighting heuristic with grid search optimization that balances personality expression with task performance. Our approach shows promise in controllably steering personality traits at the mechanistic level while maintaining high performance on standard benchmarks.

人格建模隐空间干预可解释性大模型控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。