提出超越Softmax的梯度强化学习新框架,支持动作间相关性建模。
Beyond Softmax: A New Perspective on Gradient Bandits
- 基于广义嵌套对数模型构造新算法,放宽Softmax的独立性假设
- 理论证明算法在对抗环境下可实现亚线性遗憾,涵盖Exp3等经典方法
- 适合需要处理动作相关性的强化学习场景,如推荐系统、多臂老虎机
我们建立了离散选择模型与在线学习、多臂赌博机理论之间的联系。主要贡献包括:(i) 为一类广泛算法家族提供了亚线性遗憾界,包含Exp3作为特例;(ii) 基于广义嵌套对数模型推导出新的对抗性赌博机算法;(iii) 提出一种新型广义梯度赌博机算法,突破了广泛使用的Softmax形式限制。通过放松Softmax固有的独立性假设,该框架能捕捉动作间的相关学习动态,拓展了梯度赌博机方法的应用范围。所提算法在保持闭式采样概率计算效率的同时,具备灵活的模型设定能力。数值实验在随机赌博机设置中验证了其实际有效性。
原文摘要 · Abstract (English)
We establish a link between a class of discrete choice models and the theory of online learning and multi-armed bandits. Our contributions are: (i) sublinear regret bounds for a broad algorithmic family, encompassing Exp3 as a special case; (ii) a new class of adversarial bandit algorithms derived from generalized nested logit models \citep{wen:2001}; and (iii) \textcolor{black}{we introduce a novel class of generalized gradient bandit algorithms that extends beyond the widely used softmax formulation. By relaxing the restrictive independence assumptions inherent in softmax, our framework accommodates correlated learning dynamics across actions, thereby broadening the applicability of gradient bandit methods.} Overall, the proposed algorithms combine flexible model specification with computational efficiency via closed-form sampling probabilities. Numerical experiments in stochastic bandit settings demonstrate their practical effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。