arXiv:2509.12934cs.AI2025-09被引 2

通过可解释的稀疏特征操控,揭示对齐过程中模型真实学习的内容。

The Anatomy of Alignment: Decomposing Preference Optimization by Steering Sparse Features

  • 用轻量适配器调节可解释的稀疏特征来控制模型行为
  • 发现模型更依赖风格特征而非诚实等对齐概念
  • 适合关注对齐机制透明性与可诊断性的研究者

主流对齐方法导致参数变化不透明,难以理解模型真正学到什么。为此,我们提出基于强化学习的特征操控(FSRL)框架,通过训练轻量适配器来调节可解释的稀疏特征以操控模型行为。首先,理论上证明该机制足以近似后训练过程中的行为转变。随后将FSRL应用于偏好优化,并对学习到的策略进行因果分析。结果揭示关键洞见:模型将风格表现作为质量代理,过度依赖风格和格式相关特征,而非诚实等对齐概念。通过有效优化偏好目标,FSRL成为观察对齐过程的透明代理。总体而言,FSRL提供了一个可解释的控制接口,以及在特征层面诊断偏好优化影响的实际方法。

原文摘要 · Abstract (English)

Prevailing alignment methods induce opaque parameter changes, obscuring what models truly learn. To address this, we introduce Feature Steering with Reinforcement Learning (FSRL), a framework that trains a lightweight adapter to steer model behavior by modulating interpretable sparse features. First, we theoretically demonstrate that this mechanism is expressive enough to approximate the behavioral shifts of post-training processes. We then apply FSRL to preference optimization and perform a causal analysis of the learned policy. Our analysis reveals a crucial insight: the model learns to reward stylistic presentation as a proxy for quality, disproportionately relying on features related to style and formatting over those tied to alignment concepts like honesty. By effectively optimizing the preference objective, FSRL serves as a transparent proxy for observing the alignment process. Overall, FSRL offers an interpretable control interface and a practical way to diagnose how preference optimization pressures manifest at the feature level.

对齐机制特征可解释性偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。