提出MOCHA算法,实现多目标强化学习中帕累托平稳解的高效探索。
Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach
- 融合加权切比雪夫与演员-评论家框架,系统化探索帕累托平稳解。
- 理论证明在ε-帕累托平稳解下样本复杂度为˜O(ε⁻²),依赖权重最小值p_min。
- 在KuaiRand数据集上显著优于现有基线方法,适合多目标决策场景。
在许多多目标强化学习(MORL)应用中,如何在多个非凸奖励目标下系统地探索帕累托平稳解,并具备有限时间样本复杂度保证,是一个重要但尚未充分研究的问题。为此,本文首次提出一种多目标加权切比雪夫演员-评论家(MOCHA)算法,巧妙结合加权切比雪夫(WC)与演员-评论家框架,实现帕累托平稳解的系统性探索,并提供有限时间样本复杂度保证。MOCHA算法的样本复杂度分析揭示了其对给定权重向量最小分量 $p_{ ext{min}}$ 的依赖关系,在适当选择学习率的情况下,每次探索的样本复杂度可达到 ˜𝒪(ε⁻²)。此外,在大规模 KuaiRand 离线数据集上的仿真结果表明,MOCHA算法性能显著优于其他基线MORL方法。
原文摘要 · Abstract (English)
In many multi-objective reinforcement learning (MORL) applications, being able to systematically explore the Pareto-stationary solutions under multiple non-convex reward objectives with theoretical finite-time sample complexity guarantee is an important and yet under-explored problem. This motivates us to take the first step and fill the important gap in MORL. Specifically, in this paper, we propose a \uline{M}ulti-\uline{O}bjective weighted-\uline{CH}ebyshev \uline{A}ctor-critic (MOCHA) algorithm for MORL, which judiciously integrates the weighted-Chebychev (WC) and actor-critic framework to enable Pareto-stationarity exploration systematically with finite-time sample complexity guarantee. Sample complexity result of MOCHA algorithm reveals an interesting dependency on $p_{\min}$ in finding an $ε$-Pareto-stationary solution, where $p_{\min}$ denotes the minimum entry of a given weight vector $\mathbf{p}$ in WC-scarlarization. By carefully choosing learning rates, the sample complexity for each exploration can be $\tilde{\mathcal{O}}(ε^{-2})$. Furthermore, simulation studies on a large KuaiRand offline dataset, show that the performance of MOCHA algorithm significantly outperforms other baseline MORL approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。