提出新算法解决非平稳伯努利老虎机的动态决策问题
Partition Tree Weighting for Non-Stationary Stochastic Bandits
- 将源编码中的分治树加权法拓展至带动作控制场景
- 在非平稳环境下实现高精度的即时回报预测与策略优化
- 适合研究自适应强化学习与在线学习的科研人员
本文研究交互数据流的通用源编码问题,即动作与观测交错的数据序列。目标是构建一个既具备通用性又能作为控制策略的编码分布。由于需区分动作与观测,若处理不当会引发通用设定下的自我欺骗问题。本文以具有挑战性的非平稳随机伯努利老虎机问题为场景,提出一种高效且性能优越的算法,将被动预测中的分治树加权技术推广到控制领域。
原文摘要 · Abstract (English)
This paper considers a generalisation of universal source coding for interaction data, namely data streams that have actions interleaved with observations. Our goal will be to construct a coding distribution that is both universal \emph{and} can be used as a control policy. Allowing for action generation needs careful treatment, as naive approaches which do not distinguish between actions and observations run into the self-delusion problem in universal settings. We showcase our perspective in the context of the challenging non-stationary stochastic Bernoulli bandit problem. Our main contribution is an efficient and high performing algorithm for this problem that generalises the Partition Tree Weighting universal source coding technique for passive prediction to the control setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。