用加权Lp范数改进强化学习的误差传播控制
Contraction-Aligned Analysis of Soft Bellman Residual Minimization with Weighted Lp-Norm for Markov Decision Problem
- 引入加权Lp范数软化贝尔曼残差最小化
- p越大,优化目标越贴近贝尔曼算子收缩几何
- 理论证明可更好控制误差传播,适合梯度优化
在函数逼近下求解马尔可夫决策过程仍是基本挑战,即使在线性函数逼近设定中亦然。主要难点在于几何不匹配:贝尔曼最优算子在L∞范数下具有压缩性,但常用的投影值迭代和贝尔曼残差最小化依赖于基于L2的公式。为支持梯度优化,我们考虑贝尔曼残差最小化的软化形式,并将其推广至广义加权Lp范数。我们证明,随着p增大,该公式使优化目标与贝尔曼算子的压缩几何对齐,并推导出相应的性能误差界。本分析建立了残差最小化与贝尔曼压缩之间的原理性联系,实现了对误差传播的更好控制,同时保持与梯度优化的兼容性。
原文摘要 · Abstract (English)
The problem of solving Markov decision processes under function approximation remains a fundamental challenge, even under linear function approximation settings. A key difficulty arises from a geometric mismatch: while the Bellman optimality operator is contractive in the Linfty-norm, commonly used objectives such as projected value iteration and Bellman residual minimization rely on L2-based formulations. To enable gradient-based optimization, we consider a soft formulation of Bellman residual minimization and extend it to a generalized weighted Lp -norm. We show that this formulation aligns the optimization objective with the contraction geometry of the Bellman operator as p increases, and derive corresponding performance error bounds. Our analysis provides a principled connection between residual minimization and Bellman contraction, leading to improved control of error propagation while remaining compatible with gradient-based optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。