揭示强化学习人类反馈中不同目标函数的本质等价性
When Are Two RLHF Objectives the Same?
- 提出归一化算法Opal判断两个偏好目标是否代数等价
- 发现多数常用方法实际优化同一底层目标
- 适合研究强化学习目标设计的学者参考
偏好优化文献中存在多种提出的优化目标,常被视作独立改进。本文提出Opal算法,通过生成标准形式或非等价证据来判断两个偏好目标是否代数等价。应用Opal发现,许多广泛使用的方法实际上优化相同的底层目标,而另一些则被证明确实不同。例如,批量归一化会导致同一响应对在不同批次组成下获得不同梯度。我们识别出少量结构性机制导致真正不同的目标;其余差异多为参数重写。该研究为理解偏好优化目标本质提供了新视角。
原文摘要 · Abstract (English)
The preference optimization literature contains many proposed objectives, often presented as distinct improvements. We introduce Opal, a canonicalization algorithm that determines whether two preference objectives are algebraically equivalent by producing either a canonical form or a concrete witness of non-equivalence. Applying Opal reveals that many widely used methods optimize the same underlying objective, while others are provably distinct. For example, batch normalization can cause the same response pair to receive different gradients depending on batch composition. We identify a small set of structural mechanisms that give rise to genuinely different objectives; most remaining differences are reparameterizations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。