通过方向与幅度视角改进大模型激活编辑,提升安全性表现。
Householder Pseudo-Rotation: A Novel Approach to Activation Editing in LLMs with Direction-Magnitude Perspective
- 将激活视为方向与幅度,用拟旋转操作修改
- 在多个安全评测中表现优于现有方法
- 适合关注模型行为可控性的研究者
激活编辑旨在直接修改大语言模型(LLMs)的内部表示以改变其行为并实现期望属性,已成为有前景的研究方向。现有方法主要将LLMs的激活视为空间中的点,并通过添加引导向量进行修改。然而,这种方法在提升性能的同时难以保持激活幅度的一致性。为克服这一问题,我们提出一种新编辑方法,从方向与幅度角度看待激活。所提方法名为Householder伪旋转(HPR),模拟旋转变换,从而保持激活范数,在多个安全基准测试中实现性能提升。
原文摘要 · Abstract (English)
Activation Editing, which involves directly editting the internal representations of large language models (LLMs) to alter their behaviors and achieve desired properties, has emerged as a promising area of research. Existing works primarily treat LLMs' activations as points in space and modify them by adding steering vectors. However, this approach is limited in its ability to achieve greater performance improvement while maintaining the necessary consistency of activation magnitudes. To overcome these issues, we propose a novel editing method that views activations in terms of their directions and magnitudes. Our method, named Householder Pseudo-Rotation (HPR), mimics the rotation transformation, thus preserving activation norms and resulting in an improved performance on various safety benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。