arXiv:2507.01368cs.CVcs.LG2025-07ACL被引 7

用激活值控制生成奖励信号,少样本下高效对齐模型偏好。

Activation Reward Models for Few-Shot Model Alignment

  • 通过激活值调节构建奖励信号,无需微调模型
  • 在少样本场景下优于GPT-4o等主流方法
  • 能有效防止奖励黑客行为,适合安全敏感应用

大语言模型(LLMs)和多模态模型(LMMs)与人类偏好的对齐是提升生成质量的关键挑战。传统奖励建模需独立训练奖励模型,依赖大规模偏好数据集,难以适应新偏好。为此,本文提出激活奖励模型(Activation RMs),一种少样本奖励建模新方法,利用激活值调控构建对齐奖励信号,仅需极少标注且无需额外微调。实验表明,Activation RMs在标准奖励建模基准上超越LLM-as-a-judge、基于投票评分和词元概率评分等方法。进一步提出新基准PreferenceHack,首次在成对偏好格式下测试奖励模型的抗奖励黑客能力。结果表明,Activation RMs在该基准上表现最优,优于GPT-4o。

原文摘要 · Abstract (English)

Aligning Large Language Models (LLMs) and Large Multimodal Models (LMMs) to human preferences is a central challenge in improving the quality of the models' generative outputs for real-world applications. A common approach is to use reward modeling to encode preferences, enabling alignment via post-training using reinforcement learning. However, traditional reward modeling is not easily adaptable to new preferences because it requires a separate reward model, commonly trained on large preference datasets. To address this, we introduce Activation Reward Models (Activation RMs) -- a novel few-shot reward modeling method that leverages activation steering to construct well-aligned reward signals using minimal supervision and no additional model finetuning. Activation RMs outperform existing few-shot reward modeling approaches such as LLM-as-a-judge with in-context learning, voting-based scoring, and token probability scoring on standard reward modeling benchmarks. Furthermore, we demonstrate the effectiveness of Activation RMs in mitigating reward hacking behaviors, highlighting their utility for safety-critical applications. Toward this end, we propose PreferenceHack, a novel few-shot setting benchmark, the first to test reward models on reward hacking in a paired preference format. Finally, we show that Activation RM achieves state-of-the-art performance on this benchmark, surpassing even GPT-4o.

奖励建模少样本学习对齐安全激活值控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。