arXiv:2605.12176cs.LG2026-05

多任务学习中用低秩表示提升安全线性老虎机的决策效率

Multi-Task Representation Learning for Conservative Linear Bandits

  • 通过低秩结构共享多任务特征,降低学习复杂度
  • 在安全约束下实现近似最优的累计损失(遗憾)
  • 适合需保证安全性的工业级强化学习场景

本文提出针对线性老虎机的约束型多任务表示学习框架(CMTRL)。考虑T个在d维空间中的线性带宽任务,这些任务共享一个维度为r的低维公共表示,且满足安全或性能约束,即仅允许符合特定安全要求的动作。我们设计了新型算法Safe-AltGDmin,用于在满足约束条件下恢复低秩特征矩阵。基于该算法,构建了保守线性老虎机的多任务表示学习框架,并建立了其遗憾与样本复杂度的理论保证。实验表明,该算法在多个基准方法上表现更优。

原文摘要 · Abstract (English)

This paper presents the Constrained Multi-Task Representation Learning (CMTRL) framework for linear bandits. We consider T linear bandit tasks in a d dimensional space, which share a common low-dimensional representation of dimension r, where r is much smaller than the minimum of d and T. Furthermore, tasks are constrained so that only actions meeting specific safety or performance requirements are allowed, referred to as conservative (safe) bandits. We introduce a novel algorithm, Safe-Alternating projected Gradient Descent and minimization (Safe-AltGDmin), to recover a low-rank feature matrix while satisfying the given constraints. Building on this algorithm, we propose a multi-task representation learning framework for conservative linear bandits and establish theoretical guarantees for its regret and sample complexity bounds. We presented experiments and compared the performance of our algorithm with benchmark algorithms.

多任务学习线性老虎机安全强化学习低秩表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。