arXiv:2507.11316cs.CLcs.AI2025-07ACL被引 15

通过激活内部价值向量,让大模型稳定遵循人类价值观。

Internal Value Alignment in Large Language Models through Controlled Value Vector Activation

  • 直接分析模型隐空间中的价值编码,精准定位并调控相关激活。
  • 在10种基本价值观上控制成功率最高,且不降低模型流畅度。
  • 适合需要可控伦理输出的AI系统开发者使用。

让大型语言模型(LLMs)与人类价值观对齐已受到越来越多关注,因其能提供清晰性、透明性,并适应不断变化的场景。本文提出一种受控价值向量激活方法(ConVA),通过解析价值在模型隐表示中的编码方式,修改相关激活以确保LLM内部价值的一致性。为实现准确且无偏的解读,我们提出上下文控制的价值向量识别方法;为在不损害模型性能的前提下持续控制价值,引入门控式价值向量激活机制,实现高效且最小程度的控制。实验表明,该方法在10种基础价值观上的控制成功率最高,同时保持模型性能与流畅性,并能在对抗性或恶意输入下仍确保目标价值观。源代码和数据见 https://github.com/hr-jin/ConVA。

原文摘要 · Abstract (English)

Aligning Large Language Models (LLMs) with human values has attracted increasing attention since it provides clarity, transparency, and the ability to adapt to evolving scenarios. In this paper, we introduce a Controlled Value Vector Activation (ConVA) method that directly aligns the internal values of LLMs by interpreting how a value is encoded in their latent representations and modifies relevant activations to ensure consistent values in LLMs. To ensure an accurate and unbiased interpretation, we propose a context-controlled value vector identification method. To consistently control values without sacrificing model performance, we introduce a gated value vector activation method for effective and minimum degree of value control. Experiments show that our method achieves the highest control success rate across 10 basic values without hurting LLM performance and fluency, and ensures target values even with opposite and potentially malicious input prompts. Source code and data are available at~ https://github.com/hr-jin/ConVA.

价值对齐大模型可控生成隐空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。