通过行为与神经表征对齐,精准调控大模型的风险偏好。
Steering Risk Preferences in Large Language Models by Aligning Behavioral and Neural Representations
- 用马尔可夫链蒙特卡洛方法获取行为表征,与神经表征对齐找导向向量。
- 成功在不微调的情况下,可靠改变模型风险相关输出。
- 适合研究模型可控性、伦理对齐的学者使用。
调整大型语言模型(LLM)的行为可通过修改Transformer的残差流实现,只需构造合适的“引导向量”。这种表征工程方式能有效且精准地影响模型行为,无需重新训练或微调。但如何系统性地发现这些引导向量?本文提出一种原理性方法:将基于行为的方法(特别是基于大模型的马尔可夫链蒙特卡洛)生成的隐含表征,与对应的神经表征对齐,从而提取引导向量。为验证该方法,我们聚焦于从大模型中提取隐含风险偏好,并利用对齐后的表示作为引导向量来调节其风险相关输出。实验表明,所得引导向量能有效且可靠地使模型输出符合目标行为。
原文摘要 · Abstract (English)
Changing the behavior of large language models (LLMs) can be as straightforward as editing the Transformer's residual streams using appropriately constructed "steering vectors." These modifications to internal neural activations, a form of representation engineering, offer an effective and targeted means of influencing model behavior without retraining or fine-tuning the model. But how can such steering vectors be systematically identified? We propose a principled approach for uncovering steering vectors by aligning latent representations elicited through behavioral methods (specifically, Markov chain Monte Carlo with LLMs) with their neural counterparts. To evaluate this approach, we focus on extracting latent risk preferences from LLMs and steering their risk-related outputs using the aligned representations as steering vectors. We show that the resulting steering vectors successfully and reliably modulate LLM outputs in line with the targeted behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。