提出新方法同时最大化收益并降低破产风险,解决带预算约束的强化学习难题。
Survival Multiarmed Bandits with Bootstrapping Methods
- 用历史奖励数据做自助采样估计动作价值
- 在实验中优于已有基准策略,显著降低破产概率
- 适合需要长期生存的高风险决策场景
多臂老虎机(MAB)问题被广泛研究,并已应用于多个领域。生存多臂老虎机(S-MAB)是一个开放性问题,其特点是代理必须在与观测奖励直接相关的预算内行动。由于预算耗尽会导致失败,代理的目标是既最大化期望累计奖励,又最小化破产概率。本文提出一个框架,通过引入破产规避项平衡目标函数来实现这一双重目标。动作值通过一种新颖的方法估计:从先前观测到的奖励中进行自助采样。数值实验表明,所提出的策略优于文献中的基准方法。
原文摘要 · Abstract (English)
The Multiarmed Bandits (MAB) problem has been extensively studied and has seen many practical applications in a variety of fields. The Survival Multiarmed Bandits (S-MAB) open problem is an extension which constrains an agent to a budget that is directly related to observed rewards. As budget depletion leads to ruin, an agent's objective is to both maximize expected cumulative rewards and minimize the probability of ruin. This paper presents a framework that addresses such a dual goal using an objective function balanced by a ruin aversion component. Action values are estimated through a novel approach which consists of bootstrapping samples from previously observed rewards. In numerical experiments, the policies we present outperform benchmarks from the literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。