验证大模型控制向量是否反映人类价值观的几何结构
Steering Geometry: Validating Human Value Geometry in LLM Steering Space

- 用人类价值观理论构建26000样本基准测试
- 分布驱动方法与理论预测值相关性达0.51(p<1e-13)
- 模型规模越大越接近理论结构,适合价值观对齐研究
随着大语言模型在对齐敏感场景中的广泛应用,激活控制已成为无需微调的轻量级推理时行为调控方法。然而现有工作多仅验证单一行为,难以判断控制向量是否蕴含连贯语义结构。本文基于施瓦茨基本人类价值观理论,构建包含20种价值观的26,000样本基准,评估多种分布驱动方法(如CAA、SphericalSteer、ODESteer)与行为中心方法(如COLD-Steer、BiPO)在不同模型家族和规模下的表现。结果发现:分布驱动方法能恢复与理论预测一致的价值拓扑结构(斯皮尔曼相关系数最高达0.51,p < 10⁻¹³);而行为中心方法虽有相似控制效果,但与预期价值几何无显著关联。几何保真度随模型规模提升,但在指令微调后下降。此外,几何对齐程度更高的方法,在跨价值观迁移中更符合人类一致性:正确控制某一价值可提升相容价值并抑制对立价值。代码与数据已开源。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuning methods (e.g., RLHF, DPO) for behavioral control. However, existing work typically validates steering on isolated behaviors, leaving it unclear whether steering vectors encode coherent semantic structure or merely exploit behavior-specific shortcuts. We investigate whether the latent geometry of LLM steering vectors reflects theory-specified structure in human values and morality. Using Schwartz's Theory of Basic Human Values as our primary fine-grained framework, we introduce a 26K-sample benchmark covering 20 human values and analyze distribution-driven methods (e.g., CAA, SphericalSteer, ODESteer) and behavior-centric approaches (e.g., COLD-Steer, BiPO) across diverse model families and sizes. We find that distribution-driven methods recover human value topologies aligned with theoretical predictions (Spearman $\rho$ up to 0.51, $p < 10^{-13}$). In contrast, behavior-centric methods achieve comparable steering performance but show little correlation with the expected value geometry. Geometric fidelity improves with model scale but drops after instruction tuning. Finally, better geometric alignment also leads to more human-consistent transfer across values: steering one value correctly lifts compatible values and suppresses opposing ones. Code and data are available at: https://github.com/DeepRCL/Steering_Geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。