arXiv:2511.02944cs.LGcs.AI2025-11被引 1

针对行为适应与恢复动态,提出可平衡个性化与群体效应的强化学习算法。

Power Constrained Nonstationary Bandits with Habituation and Recovery Dynamics

  • 基于罗格框架设计贪心采样算法ROGUE-TS,实现次线性后悔率。
  • 引入概率截断机制,在保持低后悔的同时保障最小探索概率。
  • 适用于微随机试验,兼顾个体适应性与群体效应检测力。

决策者常需在奖励未知且随历史策略变化的环境中选择行动,如重复使用会降低效果(习惯化),而暂停则可能恢复(恢复)。此类非平稳性由减少或增加未知效力(ROGUE)带宽框架建模,适用于行为健康干预等场景。现有算法虽能提供次线性后悔策略,但过度强调利用导致探索不足,影响对群体效应的估计。这在微随机试验(MRT)中尤为关键,其旨在开发具有群体效应的即时自适应干预并提供个性化推荐。本文首先提出针对ROGUE框架的罗格-贪心采样(ROGUE-TS)算法,并给出次线性后悔的理论保证。随后引入概率截断方法,在个人化与群体学习间实现量化权衡,平衡后悔与最低探索概率。在两项关于体力活动促进和双相情感障碍治疗的MRT数据集上验证表明,所提方法在更低后悔下仍保持高统计功效,且未显著增加后悔。该方法支持可靠检测治疗效应,同时考虑个体行为动态。对设计MRT的研究者而言,本框架提供了平衡个性化与统计有效性的实用指导。

原文摘要 · Abstract (English)

A common challenge for decision makers is selecting actions whose rewards are unknown and evolve over time based on prior policies. For instance, repeated use may reduce an action's effectiveness (habituation), while inactivity may restore it (recovery). These nonstationarities are captured by the Reducing or Gaining Unknown Efficacy (ROGUE) bandit framework, which models real-world settings such as behavioral health interventions. While existing algorithms can compute sublinear regret policies to optimize these settings, they may not provide sufficient exploration due to overemphasis on exploitation, limiting the ability to estimate population-level effects. This is a challenge of particular interest in micro-randomized trials (MRTs) that aid researchers in developing just-in-time adaptive interventions that have population-level effects while still providing personalized recommendations to individuals. In this paper, we first develop ROGUE-TS, a Thompson Sampling algorithm tailored to the ROGUE framework, and provide theoretical guarantees of sublinear regret. We then introduce a probability clipping procedure to balance personalization and population-level learning, with quantified trade-off that balances regret and minimum exploration probability. Validation on two MRT datasets concerning physical activity promotion and bipolar disorder treatment shows that our methods both achieve lower regret than existing approaches and maintain high statistical power through the clipping procedure without significantly increasing regret. This enables reliable detection of treatment effects while accounting for individual behavioral dynamics. For researchers designing MRTs, our framework offers practical guidance on balancing personalization with statistical validity.

强化学习非平稳行为健康微随机试验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。