arXiv:2511.22104cs.LGmath.OC2025-11

从海量动作中选出代表性子集,让强化学习高效决策。

Representative Action Selection for Large Action Space: From Bandits to MDPs

  • 用元学习方法选动作子集,覆盖所有环境的近优解。
  • 理论证明子集性能接近全动作空间,且在非中心高斯模型下成立。
  • 适合库存管理、推荐系统等大规模决策场景。

我们研究在共享超大动作空间的一组强化学习环境中,如何选择一个固定的小型代表性动作子集——这是库存管理与推荐系统等应用中的核心挑战,因直接在全动作空间学习不可行。目标是确保该子集对每个环境都包含一个近似最优动作,从而实现无需遍历全部动作的高效学习。本文将先前针对元贝叶斯问题的成果拓展至更通用的马尔可夫决策过程(MDP)设置。我们证明现有算法在性能上可媲美使用完整动作空间。该理论保证基于一种宽松的非中心次高斯过程模型,能容纳更高的环境异质性。因此,本方法为不确定性下的大规模组合决策提供了计算与样本高效解决方案。

原文摘要 · Abstract (English)

We study the problem of selecting a small, representative action subset from an extremely large action space shared across a family of reinforcement learning (RL) environments -- a fundamental challenge in applications like inventory management and recommendation systems, where direct learning over the entire space is intractable. Our goal is to identify a fixed subset of actions that, for every environment in the family, contains a near-optimal action, thereby enabling efficient learning without exhaustively evaluating all actions. This work extends our prior results for meta-bandits to the more general setting of Markov Decision Processes (MDPs). We prove that our existing algorithm achieves performance comparable to using the full action space. This theoretical guarantee is established under a relaxed, non-centered sub-Gaussian process model, which accommodates greater environmental heterogeneity. Consequently, our approach provides a computationally and sample-efficient solution for large-scale combinatorial decision-making under uncertainty.

强化学习动作选择决策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。