剖析大模型对齐算法如何改变内部表示,揭示其机制差异。
Mechanistic Analysis of Alignment Algorithms in Language Models

- 通过层间探针与稀疏自编码器定位偏好表征位置。
- 不同算法导致可分性提升或下降,几何变化差异显著。
- 适合研究模型可解释性与安全性的研究人员参考。
后训练对齐算法普遍被视为黑箱,掩盖了其如何重塑语言模型内部计算过程。本文系统分析六种偏好优化方法(PPO、DPO、SimPO、ORPO、GRPO、KTO)在三个开源模型家族中的表现。结合逐层线性探针、稀疏自编码器和交叉编码器,我们定位了偏好表征并量化了潜在空间中的几何变换。结果发现,偏好信号稳定集中在早期-中期或中期-晚期层,但不同目标引发定性不同的表征迁移。KTO与GRPO通过构造性特征共享及稀疏高显著性特征招募提升线性可分性;而DPO与ORPO则通过非构造性几何旋转与特征衰减降低可分性,PPO与SimPO基本保持原有几何结构。这些变换具有架构依赖性,表明行为对齐不等于统一的内部重构。研究揭示对齐是异质干预,推动以特征级审计保障安全与可解释性,并强调需设计机制感知的优化目标。
原文摘要 · Abstract (English)
Post-training alignment algorithms are predominantly evaluated as black boxes, obscuring how they reshape language models' internal computations. We present a systematic mechanistic analysis of six preference-optimization methods: PPO, DPO, SimPO, ORPO, GRPO, and KTO across three open-weight model families. By integrating layer-wise linear probing, Sparse Autoencoders, and crosscoders, we localize preference representations and quantify alignment-induced geometric transformations in latent space. We find that preference signals consistently concentrate in early--mid or mid--late layers, but different objectives induce qualitatively distinct representational shifts. KTO and GRPO enhance linear separability through constructive feature sharing and sparse, high-salience recruitment. In contrast, DPO and ORPO degrade separability via non-constructive geometric rotation and feature attenuation, while PPO and SimPO largely preserve baseline geometry. These transformations exhibit architecture-dependent variability, demonstrating that behavioral alignment does not imply uniform internal restructuring. Our findings establish alignment as a heterogeneous intervention, motivate standardized feature-level auditing for safety and interpretability, and highlight the need for mechanism-aware optimization objectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。