发现偏好训练的更新有头尾结构,头负责主要行为变化
Preference Tuning as Spectral Update Reorganization
- 将偏好更新分解为谱成分,可插拔重组干预
- 早期出现紧凑的谱头,主导最终行为变化
- 适合研究对齐机制或模型偏差的学者
基于偏好的后训练通常通过最终行为理解,但其产生的参数更新仍不透明。我们通过分析诱导参数更新的谱结构,研究了RLHF及相关偏好优化。通过分解有效LoRA更新并重新加载其谱成分作为即插即用模块,使偏好引起的更新成为可隔离、可重组和直接干预的对象。在不同模型家族、优化算法和监督模式下,这些更新均呈现出谱头-尾组织:早期出现紧凑的谱头,主导最终行为偏移,而异质性残差尾部持续存在。该分拆具有功能性而非仅描述性。插件干预表明,谱头承担可见的行为偏离,而尾部单独作用微弱。跨运行重组显示,混合适配器遵循谱头来源,说明谱头携带运行级求解器偏差。终点主导性并不意味着学习充分:仅头部学习虽非空洞但无法恢复完整解,尤其在分布外行为上;仅尾部学习几乎无可见收益,但完整解仍需尾部。这些发现将偏好后训练重构为结构化更新重组,而非单一行为修正,并表明对齐增益与覆盖损失与学习更新本身的组织方式密切相关。
原文摘要 · Abstract (English)
Preference-based post-training is usually understood through endpoint behavior, yet the learned update that produces this behavior remains largely opaque. We study RLHF and related preference optimization through the spectral structure of their induced parameter updates. By decomposing effective LoRA updates and reloading their spectral components as plug-in modules, we turn preference-induced updates into objects that can be isolated, recomposed, and directly intervened on. Across model families, optimization algorithms, and supervision regimes, these updates consistently develop a spectral head--tail organization. A compact head emerges early and carries the dominant endpoint shift, while a heterogeneous residual tail remains. The split is functional rather than merely descriptive. Plug-in intervention shows that the head accounts for the visible behavioral departure from the base model, while the tail is weak in isolation. Cross-run recomposition further shows that mixed adapters follow the source of the head, indicating that the head carries run-level solver bias. This endpoint dominance does not imply learning sufficiency. Head-only learning is non-vacuous but fails to recover the full solution, especially on out-of-distribution behavior. Tail-only learning yields little visible gain, yet the full solution is not recovered without the tail. These findings recast preference post-training as structured update reorganization rather than a monolithic behavioral correction, and suggest that alignment gain and coverage loss are tied to how the learned update itself is organized.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。