arXiv:2603.14867cs.LGcs.AI2026-03中稿 · ICAPS 2026被引 1

提出高效超梯度估计方法,解决去中心化双层强化学习难题

Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning

  • 基于Boltzmann协方差技巧,仅用交互样本即可估算超梯度
  • 在高维决策空间下仍保持高效,无需重复访问状态
  • 首次实现去中心化双人马尔可夫博弈中的超梯度优化

许多策略决策问题,如仓库机器人环境设计,可自然建模为双层强化学习(RL),其中领导者优化自身目标,而跟随者在领导者决策条件下求解马尔可夫决策过程(MDP)。当领导者无法干预跟随者的优化过程、仅能观测优化结果时,面临根本性挑战。本文通过推导领导者目标的超梯度——即反映跟随者最优策略变化的领导者策略梯度——解决此去中心化场景。不同于以往需大量数据重复访问状态或依赖复杂度随高维决策空间急剧上升的梯度估计器的方法,我们利用Boltzmann协方差技巧,提出一种替代性超梯度形式,实现仅从交互样本中高效估计超梯度,即使领导者决策空间维度很高也适用。此外,据我们所知,这是首个在去中心化设置下支持2人马尔可夫博弈超梯度优化的方法。实验验证了超梯度更新的有效性,在离散与连续状态任务中均表现优异。

原文摘要 · Abstract (English)

Many strategic decision-making problems, such as environment design for warehouse robots, can be naturally formulated as bi-level reinforcement learning (RL), where a leader agent optimizes its objective while a follower solves a Markov decision process (MDP) conditioned on the leader's decisions. In many situations, a fundamental challenge arises when the leader cannot intervene in the follower's optimization process; it can only observe the optimization outcome. We address this decentralized setting by deriving the hypergradient of the leader's objective, i.e., the gradient of the leader's strategy that accounts for changes in the follower's optimal policy. Unlike prior hypergradient-based methods that require extensive data for repeated state visits or rely on gradient estimators whose complexity can increase substantially with the high-dimensional leader's decision space, we leverage the Boltzmann covariance trick to derive an alternative hypergradient formulation. This enables efficient hypergradient estimation solely from interaction samples, even when the leader's decision space is high-dimensional. Additionally, to our knowledge, this is the first method that enables hypergradient-based optimization for 2-player Markov games in decentralized settings. Experiments highlight the impact of hypergradient updates and demonstrate our method's effectiveness in both discrete and continuous state tasks.

双层强化学习超梯度去中心化马尔可夫博弈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。