通过激活工程识别并调节大模型人格特征,实现动态性格控制。
Identifying and Manipulating Personality Traits in LLMs Through Activation Engineering
- 基于激活工程定位人格相关神经方向,实现可解释的个性调控。
- 首次在真实大模型中验证人格特征可被精准识别与修改。
- 适合关注模型可解释性与伦理安全的研究者与开发者。
近年来,大语言模型(LLMs)的发展迅速,主要目标是提升效率、可解释性与使用安全性。本文基于新兴的“激活工程”方法,探索大模型中人格特征的修改,借鉴了如《Refusal in LLMs Is Mediated by a Single Direction》(arXiv:2406.11717)和《Steering Llama 2 via Contrastive Activation Addition》(arXiv:2312.06681)等研究。我们利用激活工程,开发出一种识别并调整与人格特质相关的激活方向的方法,有望实现大模型人格的动态微调。本研究旨在深化对大模型可解释性的理解,同时探讨此类技术带来的伦理影响。
原文摘要 · Abstract (English)
The field of large language models (LLMs) has grown rapidly in recent years, driven by the desire for better efficiency, interpretability, and safe use. Building on the novel approach of "activation engineering," this study explores personality modification in LLMs, drawing inspiration from research like Refusal in LLMs Is Mediated by a Single Direction (arXiv:2406.11717) and Steering Llama 2 via Contrastive Activation Addition (arXiv:2312.06681). We leverage activation engineering to develop a method for identifying and adjusting activation directions related to personality traits, which may allow for dynamic LLM personality fine-tuning. This work aims to further our understanding of LLM interpretability while examining the ethical implications of such developments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。