arXiv:2608.07725cs.LG2026-08

提出可审计的平均奖励强化学习误差下界,提升理论比较可靠性。

Finite Constant Frontiers and Auditable Regret Certificates for Average-Reward Reinforcement Learning

  • 构建二叉树状两状态模型,精确计算轨迹级伯努利KL散度。
  • 在中等条件下误差系数达0.0200,强条件下最高0.0291,提升94%。
  • 提供可验证的误差证书,适合理论研究与算法对比分析。

平均奖励强化学习的遗憾上界已知至对数因子,但现有保证因概率模式、结构参数、对数归一化、先验信息和规划假设不同而难以比较。本文引入常数敏感的比较协议,推导出连通马尔可夫决策过程的显式有限下界。构造基于二叉树的两状态块结构,证明过程使用精确轨迹级伯努利KL散度,明确保留动作预算、直径、占据率、导航成本和终止偏差。统一闭式上界在有限前沿中将已有系数0.015提升至中等情形下的0.0200,强动作、直径与时域条件下最高达0.0291,增幅94%;极限系数为$\frac1{32}\sqrt{(A-3)/A}$。针对上界,给出一种带跨度约束的乐观学习者的可审计组合规则,但未给出具体系数;自适应方向方差与规划证书仍待解决。同时形式化了有效期望转换与常数可比性。受控诊断测试了直径依赖性、奖金与宽度交互、跨度误设,以及该下界在精确族上的表现。

原文摘要 · Abstract (English)

Average-reward reinforcement-learning regret is known up to logarithmic factors, but the numerical content of published guarantees is difficult to compare because probability mode, structural parameter, logarithmic normalization, prior information, and planning assumptions differ. We introduce a constant-aware comparison protocol and derive an explicit finite lower certificate for communicating MDPs. The construction is a binary tree of two-state blocks; its proof uses exact trajectory-level Bernoulli KL divergence and keeps action budget, diameter, occupancy, navigation cost, and terminal bias explicit. A common closed-form envelope improves the published coefficient $0.015$ across a finite frontier: $0.0200$ in a moderate regime and up to $0.0291$ under stronger action, diameter, and horizon conditions, a $94\%$ increase. The limiting coefficient is $\frac1{32}\sqrt{(A-3)/A}$. For upper bounds, we give an auditable composition rule for a span-constrained optimistic learner, but do not claim a coefficient while adaptive directional-variance and planning certificates remain open. We also formalize valid expectation conversion and constant comparability. Controlled diagnostics test diameter dependence, bonus-by-width interactions, span misspecification, and the finite lower certificate on its exact family.

强化学习后悔界理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。