arXiv:2602.23116cs.LGcs.GT2026-02被引 4

提出通用偏好下高效在线对齐新方法,突破仅限KL散度的限制。

Provably Efficient Regularized Online RLHF with Generalized Bilinear Preferences

  • 基于广义双线性偏好模型,用强凸与反对称性推导策略误差界。
  • 在特征覆盖假设下,实现与维度相关的快速收敛率,优于传统方法。
  • 适用于需要稳定对齐的强化学习场景,如智能体行为优化。

研究在一般偏好和弱反馈下的正则化在线强化学习人类对齐(RLHF)问题。尽管多种正则化被用于提升对齐鲁棒性,但现有理论中多对数后悔率仍高度依赖于KL散度。为探究此类快速率是否可推广至其他正则化,本文采用广义双线性偏好模型(GBPM),通过一个秩为2r的反对称矩阵捕捉维度为d的项目特征中的非传递偏好,以隔离通用正则化的影响。关键发现是:在GBPM下,任意贪心策略的对偶间隙由平方估计误差上界控制,仅依赖强凸性和反对称性。在特征覆盖假设下,使用贪心采样可获得通用多对数后悔率 $ ilde{ ext{O}}(ηd^4 C_{ ext{min}}^{-1} ( ext{log } T)^2 igwedge d^2 C_{ ext{min}}^{-1/2} ext{ } ext{sqrt}{T})$;而使用探索-再利用策略,在臂集条件良好时,可实现维度更优的后悔率 $ ilde{ ext{O}}(C_{ ext{min}}^{-2} ext{ } ext{sqrt}{ηr T} igwedge r^{1/3} C_{ ext{min}}^{-4/3} T^{2/3})$,其中 $η^{-1}$ 为正则化系数,$T$ 为时间跨度,$C_{ ext{min}}$ 为臂集相关量。结果表明,‘快速’后悔率并非仅源于KL,而是通用强凸几何的固有属性。

原文摘要 · Abstract (English)

We consider the problem of regularized best-response max-regret minimization in online RLHF under general preferences and bandit feedback. While various regularizers are utilized to robustify alignment, known polylogarithmic regret guarantees remain heavily specific to KL. To investigate whether such fast rates extend beyond KL, we adopt the Generalized Bilinear Preference Model (GBPM) -- capturing intransitive preferences over $d$-dimensional item-wise features via a rank-$2r$ skew-symmetric matrix -- to isolate the impact of generic regularization. Crucially, under GBPM, we prove that the dual gap of any greedy policy is bounded by the squared estimation error, derived using \emph{only} strong convexity and skew-symmetry. Under a feature coverage assumption, we establish a \emph{generic} polylogarithmic regret of $\tilde{\mathcal{O}}(ηd^4 C_{\min}^{-1} (\log T)^2 \wedge d^2 C_{\min}^{-1/2} \sqrt{T})$ with Greedy Sampling, and a dimension-wise improved regret (for well-conditioned arm-sets) of $\tilde{\mathcal{O}}(C_{\min}^{-2} \sqrt{ηr T} \wedge r^{1/3} C_{\min}^{-4/3} T^{2/3})$ with Explore-Then-Commit, where $η^{-1}$ is the regularization coefficient, $T$ is the time horizon, and $C_{\min}$ is an arm-set dependent quantity. This demonstrates that ``fast'' regrets are not KL-specific, but rather a fundamental consequence of generic strongly convex geometry.

在线学习强化学习对齐后悔率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。