提出可应对环境突变的多臂赌博机模型,让推荐系统更智能。
MARBLE: Multi-Armed Restless Bandits in Latent Markovian Environment
- 引入隐马尔可夫状态模拟环境动态变化
- 新算法在环境突变下仍能收敛到最优策略
- 适合需要实时适应变化的推荐与决策场景
非平稳环境中的决策问题常被传统多臂赌博机模型忽略。本文提出MARBLE(潜在马尔可夫环境下的多臂赌博机),通过引入随时间切换的隐马尔可夫状态来模拟非平稳行为,使每个臂的演化依赖于未观测的环境状态。为此,我们提出马尔可夫平均索引性(MAI)条件,证明在该条件下,基于惠特尔指数的同步Q-learning(QWI)几乎必然收敛至最优Q函数和对应的惠特尔指数。在嵌入校准模拟器(数字孪生)的推荐系统上验证,QWI能持续适应隐状态变化并收敛至最优策略,实证支持理论结果。
原文摘要 · Abstract (English)
Restless Multi-Armed Bandits (RMABs) are powerful models for decision-making under uncertainty, yet classical formulations typically assume fixed dynamics, an assumption often violated in nonstationary environments. We introduce MARBLE (Multi-Armed Restless Bandits in a Latent Markovian Environment), which augments RMABs with a latent Markov state that induces nonstationary behavior. In MARBLE, each arm evolves according to a latent environment state that switches over time, making policy learning substantially more challenging. We further introduce the Markov-Averaged Indexability (MAI) criterion as a relaxed indexability assumption and prove that, despite unobserved regime switches, under the MAI criterion, synchronous Q-learning with Whittle Indices (QWI) converges almost surely to the optimal Q-function and the corresponding Whittle indices. We validate MARBLE on a calibrated simulator-embedded (digital twin) recommender system, where QWI consistently adapts to a shifting latent state and converges to an optimal policy, empirically corroborating our theoretical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。