arXiv:2606.16759cs.LG2026-06

通过最大熵逆强化学习,从专家行为中恢复平均奖励下的均场博弈策略。

Maximum Entropy Inverse Reinforcement Learning for Mean-Field Games with Average Reward

论文配图:Maximum Entropy Inverse Reinforcement Learning for Mean-Field Games with Average Reward
图 1 · 摘自论文原文
  • 基于占用测度框架统一处理线性与再生核希尔伯特空间奖励
  • 提出带收敛保证的梯度上升算法,恢复策略与专家行为高度吻合
  • 适用于建模大规模群体决策,如网络攻击或消费选择场景

我们研究离散时间、无限时域均场博弈(MFGs)在平均奖励准则下的逆强化学习问题。假设专家示范来自未知奖励下的平稳均场均衡,目标是通过最大因果熵原理恢复解释观察行为的策略。通过强制与专家均场项及长期特征期望一致,将两类奖励统一在占用测度框架下。对于有限维线性奖励,给出凸对偶重构形式,具有显式对数分划目标,并证明光滑性与曲率性质,支持常步长梯度下降。对于无限维再生核希尔伯特空间(RKHS)奖励,开发拉格朗日松弛,其内层最优策略由软贝尔曼方程刻画。主要挑战在于缺乏折扣因子收缩性,我们通过引入基于小化条件的次随机核,实现软贝尔曼算子的严格收缩。建立对数似然得分的弗雷歇可微性与利普希茨光滑性,导出具收敛保证的梯度上升算法。两个数值实验——恶意软件传播MFG和基于RKHS的消费者选择模型——表明恢复策略与专家行为高度一致。

原文摘要 · Abstract (English)

We study inverse reinforcement learning for discrete-time, infinite-horizon mean-field games (MFGs) under an average-reward criterion. Expert demonstrations are assumed to arise from a stationary mean-field equilibrium under an unknown reward, and the goal is to recover a policy explaining the observed behaviour via the maximum causal entropy principle. We formulate the inverse problem by enforcing consistency with the expert mean-field term and long-run feature expectations, treating two reward classes within a unified occupation-measure framework. For finite-dimensional linear rewards, we give a convex dual reformulation with an explicit log-partition objective, and prove smoothness and curvature properties justifying constant-step-size gradient descent. For infinite-dimensional RKHS rewards, we develop a Lagrangian relaxation whose inner-maximising policy is characterised by a soft Bellman equation. The main obstacle is the absence of a discount-factor contraction. We resolve this by introducing a minorisation-based sub-stochastic kernel that yields a strict contraction of the soft Bellman operator. We establish Fréchet differentiability and Lipschitz smoothness of the log-likelihood score, leading to a gradient ascent algorithm with convergence guarantees. Two numerical examples, a malware-spread MFG and an RKHS-based consumer-choice model, show that the recovered policies closely match expert behaviour.

逆强化学习均场博弈平均奖励最大熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。