用贝叶斯方法优化健康干预算法,降低用户负担并提升学习稳定性。
Optimizing Algorithms for Mobile Health Interventions with Active Querying Optimization
- 采用卡尔曼滤波式贝叶斯更新替代传统Q-learning,实时追踪不确定性。
- 在小规模环境中实现更低方差与更稳定策略,性能接近或超越原方法。
- 适合低数据、高成本测量场景,如移动健康应用的个性化干预设计。
移动健康(mHealth)中的强化学习需平衡干预效果与用户负担,尤其当状态测量(如问卷反馈)成本高却关键。现有基于行动-条件无噪声可观察马尔可夫决策过程(ACNO-MDP)的先验-后验测量(ATM)启发式方法虽解耦控制与测量,但依赖时序差分类Q-learning,在稀疏噪声环境中易不收敛。本文提出一种贝叶斯扩展的ATM,以卡尔曼滤波风格的贝叶斯更新替代标准Q-learning,保持对Q值的不确定性估计,实现更稳定且样本高效的学习。在小规模表格环境与临床驱动测试平台中评估表明:在小规模环境中,贝叶斯ATM取得可比或更优的加权回报,方差显著降低,策略行为更稳定;但在更大更复杂的现实mHealth场景中,标准与贝叶斯版本均表现不佳,暗示ATM建模假设与真实世界结构挑战不匹配。研究强调不确定性感知方法在低数据场景的价值,也凸显需发展能显式建模因果结构、连续状态及观测延迟的新型强化学习算法,同时考虑观测成本约束。
原文摘要 · Abstract (English)
Reinforcement learning in mobile health (mHealth) interventions requires balancing intervention efficacy with user burden, particularly when state measurements (for example, user surveys or feedback) are costly yet essential. The Act-Then-Measure (ATM) heuristic addresses this challenge by decoupling control and measurement actions within the Action-Contingent Noiselessly Observable Markov Decision Process (ACNO-MDP) framework. However, the standard ATM algorithm relies on a temporal-difference-inspired Q-learning method, which is prone to instability in sparse and noisy environments. In this work, we propose a Bayesian extension to ATM that replaces standard Q-learning with a Kalman filter-style Bayesian update, maintaining uncertainty-aware estimates of Q-values and enabling more stable and sample-efficient learning. We evaluate our method in both toy environments and clinically motivated testbeds. In small, tabular environments, Bayesian ATM achieves comparable or improved scalarized returns with substantially lower variance and more stable policy behavior. In contrast, in larger and more complex mHealth settings, both the standard and Bayesian ATM variants perform poorly, suggesting a mismatch between ATM's modeling assumptions and the structural challenges of real-world mHealth domains. These findings highlight the value of uncertainty-aware methods in low-data settings while underscoring the need for new RL algorithms that explicitly model causal structure, continuous states, and delayed feedback under observation cost constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。