arXiv:2603.20453cs.LG2026-03

多源不完美反馈下,强化学习可实现更优的后悔边界。

Regret Bounds for Reinforcement Learning from Multi-Source Imperfect Preferences

  • 引入累积偏差预算,统一处理多源不完美偏好
  • 后悔上界为 $\tilde{O}(\sqrt{K/M}+ω)$,随源数 $M$ 改善
  • 揭示多源反馈何时有效,以及偏差 $ω$ 的根本限制

从人类反馈中进行强化学习(RLHF)用成对轨迹偏好替代难以定义的奖励函数。现有后悔导向理论通常假设偏好标签来自单一真实目标,但在实际系统中,反馈常为多源(标注者、专家、奖励模型、启发式规则),且因主观性、能力差异及标注/建模误差导致系统性偏差。本文通过累积不完美预算研究多源不完美偏好下的回合制强化学习:每源在 $K$ 轮中的偏好概率与理想理想器的总偏差不超过 $ω$。提出统一算法,后悔上界为 $\tilde{O}(\sqrt{K/M}+ω)$,当不完美程度小时可获得依赖于 $M$(源数)的统计增益,不完美大时仍保持对 $ω$ 的可接受加性依赖。我们给出下界 $\tilde{Ω}(\max\{\sqrt{K/M},ω\})$,揭示相对于 $M$ 的最优提升和对 $ω$ 的不可避免依赖,并构造反例表明,若将不完美反馈视为一致,可能产生高达 $\tilde{Ω}(\min\{ω\sqrt{K},K\})$ 的后悔。技术上,采用适应不完美的加权比较学习、面向价值的目标转移估计以控制隐含反馈引起的分布偏移,以及子重要性采样使加权目标可分析,从而量化了多源反馈在理论上如何改善 RLHF,以及累积不完美如何从根本上限制性能。

原文摘要 · Abstract (English)

Reinforcement learning from human feedback (RLHF) replaces hard-to-specify rewards with pairwise trajectory preferences, yet regret-oriented theory often assumes that preference labels are generated consistently from a single ground-truth objective. In practical RLHF systems, however, feedback is typically \emph{multi-source} (annotators, experts, reward models, heuristics) and can exhibit systematic, persistent mismatches due to subjectivity, expertise variation, and annotation/modeling artifacts. We study episodic RL from \emph{multi-source imperfect preferences} through a cumulative imperfection budget: for each source, the total deviation of its preference probabilities from an ideal oracle is at most $ω$ over $K$ episodes. We propose a unified algorithm with regret $\tilde{O}(\sqrt{K/M}+ω)$, which exhibits a best-of-both-regimes behavior: it achieves $M$-dependent statistical gains when imperfection is small (where $M$ is the number of sources), while remaining robust with unavoidable additive dependence on $ω$ when imperfection is large. We complement this with a lower bound $\tildeΩ(\max\{\sqrt{K/M},ω\})$, which captures the best possible improvement with respect to $M$ and the unavoidable dependence on $ω$, and a counterexample showing that naïvely treating imperfect feedback as oracle-consistent can incur regret as large as $\tildeΩ(\min\{ω\sqrt{K},K\})$. Technically, our approach involves imperfection-adaptive weighted comparison learning, value-targeted transition estimation to control hidden feedback-induced distribution shift, and sub-importance sampling to keep the weighted objectives analyzable, yielding regret guarantees that quantify when multi-source feedback provably improves RLHF and how cumulative imperfection fundamentally limits it.

强化学习偏好学习后悔边界多源反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。