提出低秩奖励框架,实现多任务强化学习的可证明高效表示学习。
Provable Multi-Task Reinforcement Learning: A Representation Learning Framework with Low Rank Rewards

- 基于无奖励学习先建数据收集策略,再用其探索估计低秩奖励矩阵。
- 在非高斯、无一致性假设下仍能准确恢复奖励矩阵,样本复杂度可控。
- 适用于共享状态空间的多任务强化学习,适合需鲁棒泛化的真实场景。
多任务表示学习(MTRL)通过在相关任务间学习共享潜在表示,提升整体学习效率。本文研究多任务强化学习(RL)中的MTRL,其中多个任务具有相同的状态-动作空间和转移概率,但奖励函数不同。考虑T个线性马尔可夫决策过程(MDPs),其奖励函数与转移动态均具有维度为d的线性特征嵌入。任务间的相关性由奖励矩阵的低秩结构刻画。由于数据受策略依赖且误差随时间累积,跨任务表示学习极具挑战。本方法采用无奖励强化学习框架,首先学习一个数据收集策略,该策略指导探索以估计未知奖励矩阵。重要的是,此策略生成的数据支持精确估计,进而实现近优策略学习。不同于依赖高斯特征、非相干性条件或最优解访问的现有方法,我们提出一种在更一般特征分布下有效的低秩矩阵估计方法。理论分析表明,在放宽假设条件下仍可实现准确的低秩矩阵恢复,并刻画了表示误差与样本复杂度的关系。利用学习到的表示,我们构造出近优策略并给出后悔界。实验表明,该方法能从有限数据中有效学习鲁棒的共享表示与任务动态。
原文摘要 · Abstract (English)
Multi-task representation learning (MTRL) is an approach that learns shared latent representations across related tasks, facilitating collaborative learning that improves the overall learning efficiency. This paper studies MTRL for multi-task reinforcement learning (RL), where multiple tasks have the same state-action space and transition probabilities, but different rewards. We consider T linear Markov Decision Processes (MDPs) where the reward functions and transition dynamics admit linear feature embeddings of dimension d. The relatedness among the tasks is captured by a low-rank structure on the reward matrices. Learning shared representations across multiple RL tasks is challenging due to the complex and policy-dependent nature of data that leads to a temporal progression of error. Our approach adopts a reward-free reinforcement learning framework to first learn a data-collection policy. This policy then informs an exploration strategy for estimating the unknown reward matrices. Importantly, the data collected under this well-designed policy enable accurate estimation, which ultimately supports the learning of an near-optimal policy. Unlike existing approaches that rely on restrictive assumptions such as Gaussian features, incoherence conditions, or access to optimal solutions, we propose a low-rank matrix estimation method that operates under more general feature distributions encountered in RL settings. Theoretical analysis establishes that accurate low-rank matrix recovery is achievable under these relaxed assumptions, and we characterize the relationship between representation error and sample complexity. Leveraging the learned representation, we construct near-optimal policies and prove a regret bound. Experimental results demonstrate that our method effectively learns robust shared representations and task dynamics from finite data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。