arXiv:2508.18741cs.LG2025-08

提出贝尔曼残差最小化的新稳定分析,实现样本复杂度突破

Stability and Generalization for Bellman Residuals

  • 构建统一李雅普诺夫势函数,关联邻近数据集的随机梯度算法
  • 证明平均稳定度为 O(1/n),首次达到最优样本复杂度指数
  • 无需方差缩减或独立采样假设,适用于标准神经网络训练

离线强化学习与离线逆强化学习旨在从固定轨迹数据中恢复近似最优的价值函数或奖励模型,但现有方法仍难以保证贝尔曼一致性。贝尔曼残差最小化(BRM)作为一种有前景的解决方案,最近发现基于随机梯度下降-上升(SGDA)的全局收敛方法。然而其在离线设置下的统计行为尚未被充分研究。本文填补了这一空白:通过引入一个统一的李雅普诺夫势函数,将相邻数据集上的SGDA过程耦合,得到O(1/n)的平均论证稳定性界——使凸-凹鞍点问题的样本复杂度指数翻倍,为最佳已知结果。该稳定性常数直接导出无需方差缩减、额外正则化或独立小批量采样假设下的O(1/n)过风险界。理论适用于标准神经网络参数化及小批量SGD。

原文摘要 · Abstract (English)

Offline reinforcement learning and offline inverse reinforcement learning aim to recover near-optimal value functions or reward models from a fixed batch of logged trajectories, yet current practice still struggles to enforce Bellman consistency. Bellman residual minimization (BRM) has emerged as an attractive remedy, as a globally convergent stochastic gradient descent-ascent based method for BRM has been recently discovered. However, its statistical behavior in the offline setting remains largely unexplored. In this paper, we close this statistical gap. Our analysis introduces a single Lyapunov potential that couples SGDA runs on neighbouring datasets and yields an O(1/n) on-average argument-stability bound-doubling the best known sample-complexity exponent for convex-concave saddle problems. The same stability constant translates into the O(1/n) excess risk bound for BRM, without variance reduction, extra regularization, or restrictive independence assumptions on minibatch sampling. The results hold for standard neural-network parameterizations and minibatch SGD.

强化学习稳定性分析贝尔曼残差离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。