arXiv:2410.05527cs.LGmath.OC2024-10ICLR被引 3

用偏好反馈解决随机多臂老虎机的奖励难定义问题

DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback

  • 直接在线学习偏好信号,无需精确奖励函数
  • 理论证明可实现√(T ln T)阶次线性后悔
  • 适合奖励难以量化的真实决策场景

restless multi-armed bandits (RMAB) 广泛用于建模受约束的序列决策问题,其中每个臂的状态按马尔可夫链演化,状态转移产生标量奖励。然而,RMAB的成功高度依赖奖励信号的可用性和质量。实际中,定义精确的奖励函数往往困难甚至不可行。本文提出 Pref-RMAB 模型,引入偏好信号:决策者仅在每个决策周期观察到被激活臂之间的成对偏好反馈,而非标量奖励。偏好反馈信息量少于标量奖励,使 Pref-RMAB 更具挑战性。为此,我们提出直接在线偏好学习(DOPL)算法,高效探索未知环境,自适应地在线收集偏好数据,并直接利用偏好反馈进行决策。理论证明 DOPL 实现子线性后悔。据我们所知,这是首个在偏好反馈下保证 √(T ln T) 阶后悔的 RMAB 算法。实验结果进一步验证了 DOPL 的有效性。

原文摘要 · Abstract (English)

Restless multi-armed bandits (RMAB) has been widely used to model constrained sequential decision making problems, where the state of each restless arm evolves according to a Markov chain and each state transition generates a scalar reward. However, the success of RMAB crucially relies on the availability and quality of reward signals. Unfortunately, specifying an exact reward function in practice can be challenging and even infeasible. In this paper, we introduce Pref-RMAB, a new RMAB model in the presence of \textit{preference} signals, where the decision maker only observes pairwise preference feedback rather than scalar reward from the activated arms at each decision epoch. Preference feedback, however, arguably contains less information than the scalar reward, which makes Pref-RMAB seemingly more difficult. To address this challenge, we present a direct online preference learning (DOPL) algorithm for Pref-RMAB to efficiently explore the unknown environments, adaptively collect preference data in an online manner, and directly leverage the preference feedback for decision-makings. We prove that DOPL yields a sublinear regret. To our best knowledge, this is the first algorithm to ensure $\tilde{\mathcal{O}}(\sqrt{T\ln T})$ regret for RMAB with preference feedback. Experimental results further demonstrate the effectiveness of DOPL.

强化学习偏好学习贝叶斯优化决策系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。