用隐空间对抗正则化提升语言模型偏好优化效果
Latent Adversarial Regularization for Offline Preference Optimization
- 通过对抗机制惩罚策略与参考模型的隐层表征差异
- 在多任务多架构下均实现性能提升,且计算开销小
- 对分布偏移和噪声更鲁棒,适合高可靠性场景
从人类反馈中学习通常依赖通过词元级正则化约束策略更新。然而,语言模型的偏好优化尤为困难,因为词元空间相似性并不等同于语义或行为相似性。为此,本文提出基于隐空间正则化的语言模型偏好优化方法。引入GANPO,通过惩罚策略模型与参考模型内部表征之间的差异来实现隐空间正则化。由于隐层表征不对应显式概率密度,借鉴GAN思想采用对抗方式最小化隐空间差异。将GANPO作为正则项集成到现有离线偏好优化目标中。在多个模型架构和任务上的实验表明,隐空间正则化带来一致性能提升。进一步对比发现,相较于词元级正则化,GANPO在分布偏移和噪声下提供更稳健的结构反馈,同时保持相近下游性能,仅带来轻微计算开销。
原文摘要 · Abstract (English)
Learning from human feedback typically relies on preference optimization that constrains policy updates through token-level regularization. However, preference optimization for language models is particularly challenging because token-space similarity does not imply semantic or behavioral similarity. To address this challenge, we leverage latent-space regularization for language model preference optimization. We introduce GANPO, which achieves latent-space regularization by penalizing divergence between the internal representations of a policy model and a reference model. Given that latent representations are not associated with explicit probability densities, we adopt an adversarial approach inspired by GANs to minimize latent-space divergence. We integrate GANPO as a regularizer into existing offline preference optimization objectives. Experiments across multiple model architectures and tasks show consistent improvements from latent-space regularization. Further, by comparing GANPO-induced inferential biases with those from token-level regularization, we find that GANPO provides more robust structural feedback under distributional shift and noise while maintaining comparable downstream performance with minor computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。