研究贝叶斯强化学习在模型错误设定下的表现,揭示其长期行为规律。
Dynamic Decision-Making under Model Misspecification: A Stochastic Stability Approach
- 将后验演化建模为信念单纯形上的马尔可夫过程,分析不同信念状态
- 发现三种后验演化模式:正确集中、错误集中和持续混合,对应不同后悔率
- 适用于需鲁棒决策的结构化强化学习场景,如金融投资、医疗实验
在模型不确定环境下,动态决策至关重要,但现有老虎机与强化学习算法多依赖模型正确设定。本文研究最常用的贝叶斯强化学习算法——汤普森采样(Thompson Sampling, TS)在模型类错误设定下的行为与性能。首先,在双臂高斯老虎机中,完整分类了后验演化的动态机制,识别出三种典型状态:正确模型集中、错误模型集中、持续信念混合,其特征由统计证据方向与模型-动作映射决定。这些状态可预测极限信念、动作频率与渐近后悔。随后,将分析推广至有限模型类,构建统一的随机稳定性框架,将后验演化视为信念单纯形上的马尔可夫过程。该框架给出两类充分条件以判别遍历性与瞬态行为,并实现后验动力学的归纳降维。结果首次提供了TS在模型错误设定下的定性与几何分类,连接贝叶斯学习与演化动力学,为结构化老虎机中的稳健决策奠定基础。
原文摘要 · Abstract (English)
Dynamic decision-making under model uncertainty is central to many economic environments, yet existing bandit and reinforcement learning algorithms rely on the assumption of correct model specification. This paper studies the behavior and performance of one of the most commonly used Bayesian reinforcement learning algorithms, Thompson Sampling (TS), when the model class is misspecified. We first provide a complete dynamic classification of posterior evolution in a misspecified two-armed Gaussian bandit, identifying distinct regimes: correct model concentration, incorrect model concentration, and persistent belief mixing, characterized by the direction of statistical evidence and the model-action mapping. These regimes yield sharp predictions for limiting beliefs, action frequencies, and asymptotic regret. We then extend the analysis to a general finite model class and develop a unified stochastic stability framework that represents posterior evolution as a Markov process on the belief simplex. This approach characterizes two sufficient conditions to classify the ergodic and transient behaviors and provides inductive dimensional reductions of the posterior dynamics. Our results offer the first qualitative and geometric classification of TS under misspecification, bridging Bayesian learning with evolutionary dynamics, and also build the foundations of robust decision-making in structured bandits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。