提出信息导向采样方法,解决强化学习中探索与利用的权衡问题
Information-directed sampling for bandits: a primer
- 用信息增益衡量探索价值,动态平衡短期损失与长期收益
- 在对称和公平硬币场景下,累积后悔呈有界或对数增长
- 适合研究强化学习与信息论交叉的物理学者阅读
多臂赌博机问题为分析序列学习中探索与利用的矛盾提供了基础框架。本文探讨信息导向采样(IDS)策略,这类启发式方法通过权衡即时损失与信息增益来实现平衡。研究聚焦于两状态伯努利赌博机这一可解析模型,以严格比较启发式策略与最优策略。通过引入修正的信息度量和调节参数,将IDS框架扩展至折扣无限时域设置。考察两类具体问题:对称赌博机与一个公平硬币的情形。在对称情况下,证明IDS实现有界累积后悔;在公平硬币情形下,其后悔随时域呈对数增长,符合经典渐近下界。本工作旨在为统计物理学者提供一篇兼具教学性与深度的综述。
原文摘要 · Abstract (English)
The Multi-Armed Bandit problem provides a fundamental framework for analyzing the tension between exploration and exploitation in sequential learning. This paper explores Information Directed Sampling (IDS) policies, a class of heuristics that balance immediate regret against information gain. We focus on the tractable environment of two-state Bernoulli bandits as a minimal model to rigorously compare heuristic strategies against the optimal policy. We extend the IDS framework to the discounted infinite-horizon setting by introducing a modified information measure and a tuning parameter to modulate the decision-making behavior. We examine two specific problem classes: symmetric bandits and the scenario involving one fair coin. In the symmetric case we show that IDS achieves bounded cumulative regret, whereas in the one-fair-coin scenario the IDS policy yields a regret that scales logarithmically with the horizon, in agreement with classical asymptotic lower bounds. This work serves as a pedagogical synthesis, aiming to bridge concepts from reinforcement learning and information theory for an audience of statistical physicists.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。