用强化学习优化产妇监护设备分配,提升资源有限下的预警效率
Optimizing Vital Sign Monitoring in Resource-Constrained Maternal Care: An RL-Based Restless Bandit Approach
- 将监护设备分配建模为带新约束的非平稳多臂老虎机问题
- 仿真显示性能比最优启发式方法高4倍
- 适合医疗资源紧缺场景下的智能设备调度研究者
孕产妇死亡仍是全球重大公共卫生挑战。通过分娩后持续监测产妇生命体征的早期预警系统,有望显著降低院内死亡率。无线生命体征监测设备能实现高效连续监控,但设备稀缺使得如何高效分配成为关键问题。本文将该问题建模为一种新型非平稳多臂老虎机(RMAB)问题,识别并处理了此前未被研究的领域特异性约束,显著增加了学习与规划难度。为此,采用强化学习中的近端策略优化(PPO)算法,通过训练策略与价值函数网络来学习最优分配策略。仿真结果表明,该方法相较最优启发式基线性能提升高达4倍。
原文摘要 · Abstract (English)
Maternal mortality remains a significant global public health challenge. One promising approach to reducing maternal deaths occurring during facility-based childbirth is through early warning systems, which require the consistent monitoring of mothers' vital signs after giving birth. Wireless vital sign monitoring devices offer a labor-efficient solution for continuous monitoring, but their scarcity raises the critical question of how to allocate them most effectively. We devise an allocation algorithm for this problem by modeling it as a variant of the popular Restless Multi-Armed Bandit (RMAB) paradigm. In doing so, we identify and address novel, previously unstudied constraints unique to this domain, which render previous approaches for RMABs unsuitable and significantly increase the complexity of the learning and planning problem. To overcome these challenges, we adopt the popular Proximal Policy Optimization (PPO) algorithm from reinforcement learning to learn an allocation policy by training a policy and value function network. We demonstrate in simulations that our approach outperforms the best heuristic baseline by up to a factor of $4$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。