arXiv:2412.08463cs.LGcs.AI2024-12被引 1

用逆强化学习自动学习健康干预中的最优奖励,提升母婴健康项目效果。

IRL for Restless Multi-Armed Bandits with Applications in Maternal and Child Health

  • 基于专家目标生成大规模理想干预路径,实现政策目标向算法的转化。
  • 提出梯度优化算法WHIRL,显著提升奖励学习效率与精度。
  • 在印度真实母婴健康数据上验证,优于现有方法且计算更快。

公共卫生实践中,常需在资源有限下监测患者并最大化其处于健康状态的时间。休息式多臂老虎机(RMAB)能有效分配资源,但传统方法假设奖励函数已知,这在实际中难以实现,尤其面对个体差异巨大的大规模人群时。本文首次将逆强化学习(IRL)引入RMAB,用于学习真实场景下的期望奖励。首先,允许专家在群体层面设定目标,并据此生成大规模专家轨迹;其次,提出算法WHIRL,通过梯度更新高效准确地学习奖励函数;第三,相比现有基线,在运行时间与准确性上均表现更优。最后,在印度数千名母婴受益人的真实数据上评估,验证了WHIRL的有效性。代码已公开:https://github.com/Gjain234/WHIRL。

原文摘要 · Abstract (English)

Public health practitioners often have the goal of monitoring patients and maximizing patients' time spent in "favorable" or healthy states while being constrained to using limited resources. Restless multi-armed bandits (RMAB) are an effective model to solve this problem as they are helpful to allocate limited resources among many agents under resource constraints, where patients behave differently depending on whether they are intervened on or not. However, RMABs assume the reward function is known. This is unrealistic in many public health settings because patients face unique challenges and it is impossible for a human to know who is most deserving of any intervention at such a large scale. To address this shortcoming, this paper is the first to present the use of inverse reinforcement learning (IRL) to learn desired rewards for RMABs, and we demonstrate improved outcomes in a maternal and child health telehealth program. First we allow public health experts to specify their goals at an aggregate or population level and propose an algorithm to design expert trajectories at scale based on those goals. Second, our algorithm WHIRL uses gradient updates to optimize the objective, allowing for efficient and accurate learning of RMAB rewards. Third, we compare with existing baselines and outperform those in terms of run-time and accuracy. Finally, we evaluate and show the usefulness of WHIRL on thousands on beneficiaries from a real-world maternal and child health setting in India. We publicly release our code here: https://github.com/Gjain234/WHIRL.

逆强化学习健康干预资源分配多臂老虎机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。