arXiv:2502.00423cs.LGstat.ME2025-02被引 2

解决高维在线决策中隐性群体差异带来的不确定性问题

High-Dimensional Linear Bandits under Stochastic Latent Heterogeneity

  • 提出隐性异质性强化学习框架,显式建模随机子群归属与群体奖励函数
  • 证明强后悔率必然线性增长,普通后悔率可达最优亚线性速率
  • 适用于需要个性化策略的场景,如精准营销、医疗干预等

本文针对在线决策中随机隐性异质性的关键挑战展开研究,即个体对行动的响应不仅依赖可观测上下文,还受未观测到的随机子群影响。现有数据驱动方法多捕捉可观测异质性,却在潜变量随机变化时失效。我们提出一种隐性异质性强化学习框架,显式建模子群概率与群体特异性奖励函数,以推广目标为应用背景。所提出的分阶段EM-贪心算法在高维条件下联合学习子群概率与奖励参数,实现最优估计与分类保证。分析揭示了该决策设置下独特现象:子群随机实现引发不可消除的分类不确定性,导致对全知强代理的亚线性后悔根本不可能。我们建立了强后悔与常规后悔的匹配上界与极小极大下界,前者必然线性增长,后者达到极小极大最优亚线性率。这些发现揭示了在线决策中的基本随机障碍,并指出了通过简单策略干预与机制设计获取潜在信息的可能解法。

原文摘要 · Abstract (English)

This paper addresses the critical challenge of stochastic latent heterogeneity in online decision-making, where individuals' responses to actions vary not only with observable contexts but also with unobserved, randomly realized subgroups. Existing data-driven approaches largely capture observable heterogeneity through contextual features but fail when the sources of variation are latent and stochastic. We propose a latent heterogeneous bandit framework that explicitly models probabilistic subgroup membership and group-specific reward functions, using promotion targeting as a motivating example. Our phased EM-greedy algorithm jointly learns latent group probabilities and reward parameters in high dimensions, achieving optimal estimation and classification guarantees. Our analysis reveals a new phenomenon unique to decision-making with stochastic latent subgroups: randomness in group realizations creates irreducible classification uncertainty, making sub-linear regret against a fully informed strong oracle fundamentally impossible. We establish matching upper and minimax lower bounds for both the strong and regular regrets, corresponding, respectively, to oracles with and without access to realized group memberships. The strong regret necessarily grows linearly, while the regular regret achieves a minimax-optimal sublinear rate. These findings uncover a fundamental stochastic barrier in online decision-making and point to potential remedies through simple strategic interventions and mechanism-design-based elicitation of latent information.

强化学习在线决策隐性异质性高维带宽

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。