arXiv:2501.08950cs.LOcs.LG2025-01被引 2

提出一种新迭代法,可精准逼近不可知函数的最小不动点。

Approximating Fixpoints of Approximated Functions

  • 用带衰减因子的Mann迭代逼近不确定函数的最小不动点
  • 在单调非扩张条件下,保证收敛至最小不动点,避免陷入次优解
  • 适用于强化学习中的最优回报计算与随机系统采样分析

不动点在计算机科学中无处不在,特别是在定量语义与验证中常需考虑非负实数上高维函数的最小不动点。本文研究当函数无法精确获得,仅能通过收敛序列逼近时,如何近似其最小不动点。聚焦于单调且非扩张函数,这类函数的不动点不唯一,标准迭代可能陷入非最小不动点。本文提出一种改进的Mann迭代方法,结合衰减因子,在合适条件下可保证收敛至目标函数的最小不动点。该结果在基于模型的强化学习中具有应用价值,可推导出对马尔可夫决策过程最优期望回报的收敛性。更广泛地,当系统函数可通过采样以给定概率误差边界逼近时(如简单随机博弈),该方法可几乎必然收敛至最小不动点。

原文摘要 · Abstract (English)

Fixpoints are ubiquitous in computer science and when dealing with quantitative semantics and verification one often considers least fixpoints of (higher-dimensional) functions over the non-negative reals. We show how to approximate the least fixpoint of such functions, focusing on the case in which they are not known precisely, but represented by a sequence of approximating functions that converge to them. We concentrate on monotone and non-expansive functions, for which uniqueness of fixpoints is not guaranteed and standard fixpoint iteration schemes might get stuck at a fixpoint that is not the least. Our main contribution is the identification of an iteration scheme, a variation of Mann iteration with a dampening factor, which, under suitable conditions, is shown to guarantee convergence to the least fixpoint of the function of interest. We then argue that these results are relevant in the context of model-based reinforcement learning for Markov decision processes, showing how the proposed iteration scheme instantiates and allows us to derive convergence to the optimal expected return. More generally, we show that our results can be used to iterate to the least fixpoint almost surely for systems where the function of interest can be approximated with given probabilistic error bounds, as it happens for probabilistic systems, such as simple stochastic games, which can be explored via sampling.

不动点强化学习迭代算法概率逼近

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。