从偏好反馈中学习核函数马尔可夫决策过程,实现高效策略优化
Learning Kernel-Based MDPs from Episodic Preferential Feedback

- 基于核函数建模奖励与转移,通过轨迹偏好估计价值
- 在有限轮次内实现亚线性后悔率,策略性能逼近最优
- 适用于依赖人类偏好而非数值奖励的强化学习场景
人类反馈常以偏好形式出现而非精确的数值奖励,推动了仅凭偏好进行强化学习的研究,即人类反馈强化学习(RLHF)。本文针对广义的核函数马尔可夫决策过程(kernel MDPs)中的纯偏好学习进行了严谨的理论分析。每轮中,学习者从同一初始状态执行两条策略,获得一个二元标签,表示哪条轨迹更受偏好,该偏好由累积(未观测)奖励差值的Bradley-Terry-Luce模型刻画。在对奖励和转移函数采用核函数假设(最通用且可理论分析的模型之一)的前提下,我们设计了面向期末比较的偏好型价值估计与置信集。证明了高概率后悔界随轮次数亚线性增长,表明所学策略的价值收敛于最优策略价值。
原文摘要 · Abstract (English)
Human feedback often arrives as preferences rather than calibrated numeric rewards, motivating reinforcement learning from preferential feedback, also referred to as reinforcement learning from human feedback (RLHF). We present a rigorous theoretical study of preference-only learning in episodic kernel MDPs. In each episode, the learner deploys two policies from a common start state and receives a single binary label indicating which trajectory is preferred, modeled by a Bradley--Terry--Luce link on the difference of cumulative (unobserved) rewards. Under kernel-based assumptions on the reward and transition functions (one of the most general models amenable to theoretical analysis) we develop preference-based value estimation and confidence sets tailored to end-of-episode comparisons. We prove high-probability regret bounds that scale sublinearly in the number of episodes, implying that the value of the learned policy converges to that of the optimal policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。