解决市场信息不全且有时间相关性的决策难题。
Partially Observable Contextual Bandits with Linear Payoffs
- 用系统识别与滤波融合上下文带宽算法,迭代估计隐藏参数并做决策。
- 在滤波条件满足时,采用汤普森采样可实现次线性后悔率。
- 适合金融等含隐变量、动态相关性的实际场景应用。
标准上下文老虎机假设上下文完全可观测且可操作。本文研究一种新型老虎机设置,其上下文部分可观测且存在相关性,动机源于金融领域:决策基于具有时间相关性但未被完全观测的市场信息。本文将统计信号处理思想与老虎机结合,提出名为EMKF-Bandit的算法流程,该流程整合系统辨识、滤波与经典上下文老虎机算法,形成在隐参数估计与决策间迭代的机制。当选用汤普森采样作为老虎机算法时,分析表明在滤波条件满足下,该方法可实现次线性后悔率。通过数值模拟验证了该方法的优势与实际适用性。
原文摘要 · Abstract (English)
The standard contextual bandit framework assumes fully observable and actionable contexts. In this work, we consider a new bandit setting with partially observable, correlated contexts and linear payoffs, motivated by the applications in finance where decision making is based on market information that typically displays temporal correlation and is not fully observed. We make the following contributions marrying ideas from statistical signal processing with bandits: (i) We propose an algorithmic pipeline named EMKF-Bandit, which integrates system identification, filtering, and classic contextual bandit algorithms into an iterative method alternating between latent parameter estimation and decision making. (ii) We analyze EMKF-Bandit when we select Thompson sampling as the bandit algorithm and show that it incurs a sub-linear regret under conditions on filtering. (iii) We conduct numerical simulations that demonstrate the benefits and practical applicability of the proposed pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。