提出在线强化学习的采样复杂度分析,覆盖多种动态系统。
The Sample Complexity of Online Reinforcement Learning: A Multi-model Perspective
- 从多模型视角建模连续状态动作空间的非周期系统
- 最一般情形下策略遗憾为 $\mathcal{O}(N ε^2 + d_u\ln(m(ε))/ε^2)$
- 适用于含先验知识、行为稳定的实用算法,适合理论研究者
研究在非周期设置下,具有连续状态和动作空间的非线性动力系统的在线强化学习采样复杂度。分析涵盖从有限个非线性候选模型到有界且Lipschitz连续动力学,再到参数由紧实实值集定义的系统。在最一般情形下,算法实现策略遗憾 $\mathcal{O}(N ε^2 + d_\mathrm{u} \ln(m(ε))/ε^2)$,其中 $N$ 为时间范围,$ε$ 为用户指定的离散化宽度,$d_\mathrm{u}$ 为输入维度,$m(ε)$ 通过打包数衡量函数类复杂度。在动力学参数由紧实实值集定义的情形(如神经网络、Transformer等),证明策略遗憾为 $\mathcal{O}(\sqrt{d_\mathrm{u} N p})$,其中 $p$ 为参数数量,恢复了此前针对线性时不变系统的样本复杂度结果。虽然本文聚焦采样复杂度刻画,但所提算法因结构简单、可融入先验知识且瞬态行为良好,可能具备实际应用价值。
原文摘要 · Abstract (English)
We study the sample complexity of online reinforcement learning in the general \hzyrev{non-episodic} setting of nonlinear dynamical systems with continuous state and action spaces. Our analysis accommodates a large class of dynamical systems ranging from a finite set of nonlinear candidate models to models with bounded and Lipschitz continuous dynamics, to systems that are parametrized by a compact and real-valued set of parameters. In the most general setting, our algorithm achieves a policy regret of $\mathcal{O}(N ε^2 + d_\mathrm{u}\mathrm{ln}(m(ε))/ε^2)$, where $N$ is the time horizon, $ε$ is a user-specified discretization width, $d_\mathrm{u}$ the input dimension, and $m(ε)$ measures the complexity of the function class under consideration via its packing number. In the special case where the dynamics are parametrized by a compact and real-valued set of parameters (such as neural networks, transformers, etc.), we prove a policy regret of $\mathcal{O}(\sqrt{d_\mathrm{u}N p})$, where $p$ denotes the number of parameters, recovering earlier sample-complexity results that were derived for linear time-invariant dynamical systems. While this article focuses on characterizing sample complexity, the proposed algorithms are likely to be useful in practice, due to their simplicity, their ability to incorporate prior knowledge, and their benign transient behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。