arXiv:2605.28364stat.MLcs.LG2026-05

提出可自适应方差的强化学习算法,提升学习效率。

Variance-Adaptive Optimal Algorithm for Reinforcement Learning with Multinomial Logit Function Approximation

  • 基于多项式逻辑函数近似,设计方差自适应算法
  • 实现实例最优的后悔率,理论边界更紧
  • 适合关注学习效率与稳定性研究者

基于多项式逻辑(MNL)函数近似的强化学习因其灵活性和广泛适用性已成为重要框架。尽管现有研究在最坏情况分析下建立了后悔率保证,但未能捕捉学习者与环境交互中变异性的性能影响。本文提出了针对基于MNL的马尔可夫决策过程的新理论分析,得到显式的方差自适应后悔界。所提算法计算高效,实现了实例层面最优的后悔率,缩小了上下界差距。数值实验验证,该方法比传统方法更高效地学习最优策略。

原文摘要 · Abstract (English)

Reinforcement learning with multinomial logistic (MNL) function approximation has become an important framework due to its flexibility and broad applicability. While existing studies have established regret guarantees under worst-case analysis, they do not capture how performance depends on the variability of the interaction between the learner and the environment. In this paper, we develop a new theoretical analysis for MNL-based Markov decision processes that yields explicit variance-adaptive regret bounds. Our algorithm is computationally efficient and achieves the instance-wise optimal rate of regret, narrowing the gap between upper and lower bounds. Our numerical experiments validate that our method learns optimal policies more efficiently than conventional approaches.

强化学习后悔率算法优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。