通过温度与倾斜调整,提升生成模型对齐的鲁棒性与适应性。
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
- 引入参考模型温度调节,扩展推理时对齐方法
- 构建锐化对数意见池(SLOP)提升泛化能力
- 校准权重参数有效防住奖励欺骗,适合持续优化场景
推理时对齐技术为昂贵的强化学习提供轻量级替代或补充,支持对齐目标和奖励指标动态变化时的持续适应。现有理论分析表明这些方法可近似于从最优倾斜分布中采样。本文通过引入参考模型温度调整,将推理时对齐推广至生成式奖励模型集成的锐化对数意见池(SLOP)。为缓解奖励欺骗问题,提出一种用于校准SLOP权重参数的算法,并实验验证其在保持对齐性能的同时显著提升鲁棒性。
原文摘要 · Abstract (English)
Inference-time alignment techniques offer a lightweight alternative or complement to costly reinforcement learning, while enabling continual adaptation as alignment objectives and reward targets evolve. Existing theoretical analyses justify these methods as approximations to sampling from distributions optimally tilted toward a given reward model. We extend these techniques by introducing reference-model temperature adjustment, which leads to further generalization of inference-time alignment to ensembles of generative reward models combined as a sharpened logarithmic opinion pool (SLOP). To mitigate reward hacking, we propose an algorithm for calibrating SLOP weight parameters and experimentally demonstrate that it improves robustness while preserving alignment performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。