arXiv:2502.05974cs.LGstat.ML2025-02被引 3

提出新框架精准刻画混合环境下的决策复杂度,提升学习效率。

Decision Making in Hybrid Environments: A Model Aggregation Approach

  • 通过扩展决策估计系数,统一处理固定动态与奖励突变的混合场景。
  • 在混合环境中实现更优的后悔上界,改进线性Q*/V* MDP的已知结果。
  • 适用于模型基于与无模型学习,适合研究在线决策与强化学习者。

Foster等(2021, 2022, 2023b)及Xu和Zeevi(2023)提出了决策估计系数(DEC)框架,用于刻画一般在线决策问题的复杂度并指导算法设计。然而,现有工作或聚焦于世界恒定的纯随机场景,或聚焦于世界任意变化的纯对抗场景。对于世界动态固定但奖励任意变化的混合场景,仅给出悲观的决策复杂度上界。本文提出一种DEC的通用扩展,更精确刻画该情形。该框架不仅适用于特定情况,还支持学习者在假设集子集上学习,权衡估计复杂度与决策复杂度,具有独立意义。本文涵盖混合场景下的模型基于与无模型学习,并将Du等(2021)的双线性类扩展至对抗奖励情形。此外,所提方法改进了纯随机场景下线性Q*/V* MDP的最佳已知后悔界。

原文摘要 · Abstract (English)

Recent work by Foster et al. (2021, 2022, 2023b) and Xu and Zeevi (2023) developed the framework of decision estimation coefficient (DEC) that characterizes the complexity of general online decision making problems and provides a general algorithm design principle. These works, however, either focus on the pure stochastic regime where the world remains fixed over time, or the pure adversarial regime where the world arbitrarily changes over time. For the hybrid regime where the dynamics of the world is fixed while the reward arbitrarily changes, they only give pessimistic bounds on the decision complexity. In this work, we propose a general extension of DEC that more precisely characterizes this case. Besides applications in special cases, our framework leads to a flexible algorithm design where the learner learns over subsets of the hypothesis set, trading estimation complexity with decision complexity, which could be of independent interest. Our work covers model-based learning and model-free learning in the hybrid regime, with a newly proposed extension of the bilinear classes (Du et al., 2021) to the adversarial-reward case. In addition, our method improves the best-known regret bounds for linear Q*/V* MDPs in the pure stochastic regime.

在线决策强化学习后悔界混合环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。