arXiv:2503.00345cs.LG2025-03被引 1

首次证明非线性表示学习能提升多任务强化学习效率。

Towards Understanding the Benefit of Multitask Representation Learning in Decision Process

  • 用新算法联合学习多个任务的共享非线性表示
  • 理论证明其误差上界优于独立学每个任务
  • 适合研究强化学习表示机制或迁移学习的学者

多任务表示学习(MRL)已成为提升强化学习样本效率的主流方法。尽管实证表明在在线和迁移学习中同时训练多个任务可显著提高效率,但其理论机制仍不完善。以往分析多假设表示函数已知或为线性形式,这在现实中不成立。真实场景通常需使用神经网络等非线性函数作为表示,且这些函数需从零学习。本文将分析扩展至未知的非线性表示,考虑一个智能体同时处理 M 个上下文赌博机(或马尔可夫决策过程),通过新型广义函数置信上界算法(GFUCB)从非线性函数类 Φ 中学习共享表示函数 ϕ。我们首次在一般函数类中严格证明该方法的累计后悔上界优于分别学习 M 个任务的下界,验证了 MRL 的有效性。该框架还解释了表示在面对新相关任务时如何促进迁移学习,并识别出成功迁移的关键条件。实验结果进一步支持理论发现。

原文摘要 · Abstract (English)

Multitask Representation Learning (MRL) has emerged as a prevalent technique to improve sample efficiency in Reinforcement Learning (RL). Empirical studies have found that training agents on multiple tasks simultaneously within online and transfer learning environments can greatly improve efficiency. Despite its popularity, a comprehensive theoretical framework that elucidates its operational efficacy remains incomplete. Prior analyses have predominantly assumed that agents either possess a pre-known representation function or utilize functions from a linear class, where both are impractical. The complexity of real-world applications typically requires the use of sophisticated, non-linear functions such as neural networks as representation function, which are not pre-existing but must be learned. Our work tries to fill the gap by extending the analysis to \textit{unknown non-linear} representations, giving a comprehensive analysis for its mechanism in online and transfer learning setting. We consider the setting that an agent simultaneously playing $M$ contextual bandits (or MDPs), developing a shared representation function $ϕ$ from a non-linear function class $Φ$ using our novel Generalized Functional Upper Confidence Bound algorithm (GFUCB). We formally prove that this approach yields a regret upper bound that outperforms the lower bound associated with learning $M$ separate tasks, marking the first demonstration of MRL's efficacy in a general function class. This framework also explains the contribution of representations to transfer learning when faced with new, yet related tasks, and identifies key conditions for successful transfer. Empirical experiments further corroborate our theoretical findings.

强化学习多任务学习表示学习理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。