arXiv:2410.01249cs.LG2024-10被引 2

用对偶散度优化策略,让强化学习更快更稳收敛

Dual Approximation Policy Optimization

  • 用对偶Bregman散度替代传统误差度量进行策略更新
  • 在通用函数逼近下实现线性快速收敛
  • 统一了多种经典方法,提供强理论保证

我们提出双逼近策略优化(DAPO),将通用函数逼近融入策略镜面下降方法。与常用$ L_2 $范数度量函数逼近误差的流行方法不同,DAPO采用由镜面映射诱导的对偶Bregman散度进行策略投影。该对偶框架兼具理论与实践意义:不仅在通用函数逼近下实现快速线性收敛,还包含多种著名实用方法作为特例,从而立即提供强有力的收敛保证。

原文摘要 · Abstract (English)

We propose Dual Approximation Policy Optimization (DAPO), a framework that incorporates general function approximation into policy mirror descent methods. In contrast to the popular approach of using the $L_2$-norm to measure function approximation errors, DAPO uses the dual Bregman divergence induced by the mirror map for policy projection. This duality framework has both theoretical and practical implications: not only does it achieve fast linear convergence with general function approximation, but it also includes several well-known practical methods as special cases, immediately providing strong convergence guarantees.

强化学习策略优化收敛性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。