针对隐藏状态变化的强化学习问题,提出自适应探索算法提升决策性能。
Adaptive Exploration for Latent-State Bandits
- 用滞后动作-奖励对和探测指纹双重摘要隐状态,指导策略选择。
- 在合成测试中显著降低动态损失,优于传统及非平稳基线方法。
- 适合状态随时间演变且需实时反馈的动态环境应用。
我们研究奖励依赖于不可观测马尔可夫状态的老虎机问题,该状态独立于学习者动作演化,最优动作可能随时间改变,而学习者仅能观测历史动作与奖励。提出将LinUCB与两种隐状态摘要结合:滞后动作-奖励对,以及在可用时由多臂奖励构成的探测指纹。自适应变体通过残差、边界和过期性检测更新指纹。在状态数量、转移速率、噪声水平和时间跨度等合成压力测试中,当摘要能有效区分状态且更新足够频繁时,该方法相较标准、对抗性和非平稳基线显著降低动态遗憾。消融实验与误设测试揭示主要失效模式:指纹分离弱、噪声高、序列探测期间状态突变。
原文摘要 · Abstract (English)
We study bandits whose rewards depend on an unobserved Markov state that evolves independently of the learner's actions. The optimal arm can change even though the learner observes only past actions and rewards. We propose algorithms that feed LinUCB with two summaries of the hidden state: a lagged action-reward pair and, when available, a probe fingerprint formed from rewards of multiple arms. The adaptive variants refresh the fingerprint using residual, margin, and staleness tests. In synthetic stress tests over state count, transition rate, noise, and horizon, these methods reduce dynamic regret relative to standard, adversarial, and non-stationary bandit baselines when the summaries distinguish states and are updated often enough. Ablations and misspecification tests identify the main failure modes: weak fingerprint separation, high noise, and state changes during sequential probes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。