arXiv:2601.09825cs.LG2026-01NeurIPS被引 2

提出局部化方法,突破传统分析瓶颈,实现强化学习首次真正的一阶收益界。

Eluder dimension: localise it!

  • 引入埃尔德维维数局部化新方法,突破经典分析局限
  • 在有限时域任务中首次获得真正的第一阶后悔界,优于以往结果
  • 适用于累积回报有界的强化学习场景,适合算法优化研究者

我们建立了广义线性模型类的埃尔德维维维数下界,表明基于标准埃尔德维维维数的分析无法导出一阶后悔界。为解决此问题,我们提出埃尔德维维维数的局部化方法;该分析可立即恢复并改进伯努利老虎机的经典结果,并首次实现累积回报有界的有限时域强化学习任务中的真正一阶边界。

原文摘要 · Abstract (English)

We establish a lower bound on the eluder dimension of generalised linear model classes, showing that standard eluder dimension-based analysis cannot lead to first-order regret bounds. To address this, we introduce a localisation method for the eluder dimension; our analysis immediately recovers and improves on classic results for Bernoulli bandits, and allows for the first genuine first-order bounds for finite-horizon reinforcement learning tasks with bounded cumulative returns.

强化学习后悔界算法分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。