arXiv:2411.07618cs.AIcs.CL2024-11ICML被引 11

用稀疏自编码器实现高效稳定的语言模型对齐

Constrain Alignment with Sparse Autoencoders

  • 基于预训练稀疏自编码器的特征约束优化方法
  • 在基准数据集上胜率提升5.08%,计算成本更低
  • 适合追求高效可控对齐的模型开发者

大语言模型与人类偏好对齐仍是关键挑战。尽管基于人类反馈的强化学习(RLHF)和直接偏好优化(DPO)取得显著进展,但常伴随计算效率低和训练不稳定问题。本文提出特征级约束偏好优化(FPO),利用预训练稀疏自编码器(SAEs)引入特征级约束,实现高效且稀疏强化的对齐。该方法通过使用在良好训练的稀疏自编码器中激活的稀疏特征,结合特征级离线参考的序列KL散度,提升对齐质量。在基准数据集上的实验表明,FPO相比最先进基线实现了5.08%的绝对胜率提升,同时计算成本大幅降低,为高效、可控制的大语言模型对齐提供了有前景的解决方案。

原文摘要 · Abstract (English)

The alignment of large language models (LLMs) with human preferences remains a key challenge. While post-training techniques like Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) have achieved notable success, they often introduce computational inefficiencies and training instability. In this paper, we propose Feature-level constrained Preference Optimization (FPO), a novel method designed to simplify the alignment process while ensuring stability. FPO leverages pre-trained Sparse Autoencoders (SAEs) and introduces feature-level constraints, allowing for efficient, sparsity-enforced alignment. Our approach enjoys efficiency by using sparse features activated in a well-trained sparse autoencoder and the quality of sequential KL divergence by using the feature-level offline reference. Experimental results on benchmark datasets demonstrate that FPO achieves a 5.08% absolute improvement in win rate with much lower computational cost compared to state-of-the-art baselines, making it a promising solution for efficient and controllable LLM alignments.

模型对齐稀疏自编码器高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。