arXiv:2410.21601cs.LG2024-10NeurIPS被引 1

通过低秩结构提升迁移强化学习效率,降低复杂度瓶颈。

The Limits of Transfer Reinforcement Learning with Latent Low-rank Structure

  • 利用潜在低秩结构建模源任务,提取可迁移表征
  • 在目标任务中将后悔率依赖从SA降至仅与秩相关
  • 理论证明算法对多数情形达到最优,适合复杂环境迁移

许多强化学习算法因状态空间 $S$ 和动作空间 $A$ 过大而难以实用。为解决此问题,本文研究具有潜在低秩结构的迁移强化学习。考虑源与目标马尔可夫决策过程(MDP)的转移核具有不同形式的Tucker秩:$(S, d, A)$、$(S, S, d)$、$(d, S, A)$ 或 $(d, d, d)$。每种情形下引入可迁移系数 $α$ 衡量表征迁移难度。所提算法在源MDP中学习潜在表示,并利用其线性结构消除目标MDP后悔界对 $S$、$A$ 或 $SA$ 的依赖。同时给出信息论下界,表明除 $(d,d,d)$ 情形外,算法在 $α$ 下为极小极大最优。

原文摘要 · Abstract (English)

Many reinforcement learning (RL) algorithms are too costly to use in practice due to the large sizes $S, A$ of the problem's state and action space. To resolve this issue, we study transfer RL with latent low rank structure. We consider the problem of transferring a latent low rank representation when the source and target MDPs have transition kernels with Tucker rank $(S , d, A )$, $(S , S , d), (d, S, A )$, or $(d , d , d )$. In each setting, we introduce the transfer-ability coefficient $α$ that measures the difficulty of representational transfer. Our algorithm learns latent representations in each source MDP and then exploits the linear structure to remove the dependence on $S, A $, or $S A$ in the target MDP regret bound. We complement our positive results with information theoretic lower bounds that show our algorithms (excluding the ($d, d, d$) setting) are minimax-optimal with respect to $α$.

迁移学习强化学习低秩结构理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。