让语言模型在偏好数据中自动学习隐含奖励,无需预设函数形式。
Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model
- 基于未知链接函数的半参数建模,将偏好转化为单指标选择模型。
- 证明了不依赖链接函数的收敛性,误差与函数复杂度相关。
- 适合研究模型对齐机制或想规避假设偏差的研究者使用。
策略对齐通常假设观测偏好与潜在奖励之间的链接函数已知(如Bradley-Terry模型/逻辑链接)。若链接函数设定错误,会扭曲推断出的奖励并导致策略错位。本文研究在链接函数未知且无限制情况下的策略对齐问题。提出一个 $f$-散度约束的奖励最大化问题,证明在策略类可实现条件下,可导出半参数单指标二元选择模型:一个标量策略诱导的指数捕捉演示数据的所有依赖关系,其余偏好分布保持非参数自由。不同于计量经济学中对结构参数的可识别性假设和估计,我们直接学习策略,隐式定义奖励函数,分析与最优策略的误差,并允许不可识别和非参数化指数。证明了基于通用函数复杂度度量的链接无关收敛保证,并通过实验证明方法与理论有效性。代码见 https://github.com/causalml/spo/。
原文摘要 · Abstract (English)
Policy alignment to preference data typically assumes a known link function between observed preferences and latent rewards (e.g., Bradley-Terry model / logistic link). Misspecification of this link can bias inferred rewards and misalign learned policies. We study policy alignment under an unknown and unrestricted link function. We formulate an $f$-divergence-constrained reward maximization problem and show that realizability in a policy class induces a semiparametric single-index binary choice model, where a scalar policy-induced index captures all dependence on demonstrations and the remaining preference distribution is unrestricted. Rather than impose identifiability of structural parameters of such a model and estimate them, as in econometrics, we develop methods that directly learn policies, with the reward function implicit, analyzing error to the optimal policy and allowing for unidentifiable and nonparametric indices. We prove link-agnostic convergence guarantees in terms of generic function complexity measures and validate the methods and theory empirically. Code is available at https://github.com/causalml/spo/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。