arXiv:2410.02068cs.LGstat.ML2024-10ICML被引 11

通过低秩表示提升多任务强化学习效率,显著减少样本消耗。

Fast and Sample Efficient Multi-Task Representation Learning in Stochastic Contextual Bandits

  • 用交替投影梯度下降恢复低秩特征矩阵
  • 在T个任务中实现近似最优的累计损失(后悔值)
  • 适合高维数据下的多任务在线学习场景

我们研究了表示学习如何提升上下文随机决策问题的学习效率。考虑同时进行T个维度为d的线性上下文老虎机任务,这些任务共享一个维度仅为r(远小于d)的公共线性表示。我们提出一种基于交替投影梯度下降和最小化估计器的新算法,用于恢复低秩特征矩阵。利用该估计器,我们设计了一种多任务学习算法,并给出了该算法的后悔界。实验结果表明,所提算法在性能上优于基准方法。

原文摘要 · Abstract (English)

We study how representation learning can improve the learning efficiency of contextual bandit problems. We study the setting where we play T contextual linear bandits with dimension d simultaneously, and these T bandit tasks collectively share a common linear representation with a dimensionality of r much smaller than d. We present a new algorithm based on alternating projected gradient descent (GD) and minimization estimator to recover a low-rank feature matrix. Using the proposed estimator, we present a multi-task learning algorithm for linear contextual bandits and prove the regret bound of our algorithm. We presented experiments and compared the performance of our algorithm against benchmark algorithms.

多任务学习上下文老虎机表示学习低秩估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。