arXiv:2501.10722cs.LGstat.ML2025-01

提出统一正则化方法,解决高维张量强化学习难题

A Unified Regularization Approach to High-Dimensional Generalized Tensor Bandits

  • 用凸优化与可分解正则化建模张量低秩结构
  • 理论证明比现有方法更优且适用范围更广
  • 适合处理高维复杂上下文的决策问题

现代决策场景常面临高维且富含高阶上下文信息的数据,现有强化学习算法难以生成有效策略。本文提出一种广义线性张量强化学习算法,通过引入低维张量结构,并建立统一的分析框架。该框架采用弱可分解正则化与凸优化方法,不仅在张量低秩假设下表现更优,还可扩展至切片稀疏性、低秩等多种低维结构。理论分析表明,相比已有低秩张量结果,本框架不仅提供更紧的泛化误差界,且适用性更广。在退化为低秩矩阵的特殊情形下,其边界在某些场景仍具优势。

原文摘要 · Abstract (English)

Modern decision-making scenarios often involve data that is both high-dimensional and rich in higher-order contextual information, where existing bandits algorithms fail to generate effective policies. In response, we propose in this paper a generalized linear tensor bandits algorithm designed to tackle these challenges by incorporating low-dimensional tensor structures, and further derive a unified analytical framework of the proposed algorithm. Specifically, our framework introduces a convex optimization approach with the weakly decomposable regularizers, enabling it to not only achieve better results based on the tensor low-rankness structure assumption but also extend to cases involving other low-dimensional structures such as slice sparsity and low-rankness. The theoretical analysis shows that, compared to existing low-rankness tensor result, our framework not only provides better bounds but also has a broader applicability. Notably, in the special case of degenerating to low-rank matrices, our bounds still offer advantages in certain scenarios.

强化学习张量建模正则化高维数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。