MINTS通过简化先验设计,实现带约束的高效强化学习决策。
MINTS: Minimalist Thompson Sampling
- 仅对最优位置设先验,用轮廓似然消除干扰参数
- 在均值约束下实现近最优非渐近后悔率,适配单峰结构
- 适合需快速适应复杂约束的在线决策场景
贝叶斯范式为不确定性下的序列决策提供了严谨工具,但其对所有参数依赖概率模型的特性,限制了复杂结构约束的融入。我们提出一种极简贝叶斯框架,仅对最优位置设置先验,通过轮廓似然消除干扰参数,得到可自然容纳结构约束的广义后验。作为直接实例,我们构建了极简汤普森采样(MINTS)。针对具有均值约束的多臂赌博机,我们建立了近最优的非渐近后悔率保证和精确几乎必然的渐近后悔刻画。特别地,MINTS在无结构设定下达到经典的Lai--Robbins常数,并自动适应单峰结构,实现仅由最优臂邻近臂决定的精确常数。
原文摘要 · Abstract (English)
The Bayesian paradigm offers principled tools for sequential decision-making under uncertainty, but its reliance on a probabilistic model for all parameters can hinder the incorporation of complex structural constraints. We introduce a minimalist Bayesian framework that places a prior only on the location of the optimum, while eliminating nuisance parameters through profile likelihood. This yields a generalized posterior that naturally accommodates structural constraints. As a direct instantiation, we develop MINimalist Thompson Sampling (MINTS). For multi-armed bandits with mean constraints, we establish near-optimal non-asymptotic regret guarantees and sharp almost-sure asymptotic regret characterizations. In particular, MINTS attains the classical Lai--Robbins constant in the unstructured setting and automatically adapts to unimodal structure, achieving the sharp constant determined only by the immediate neighbors of the optimal arm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。