arXiv:2507.14529cs.LGmath.OC2025-07被引 3

用核方法提升平均场博弈中奖励函数的推断能力,实现更精准的行为恢复。

Kernel Based Maximum Entropy Inverse Reinforcement Learning for Mean-Field Games

  • 在再生核希尔伯特空间中建模未知奖励,支持非线性结构直接学习。
  • 在交通路由场景中,误差比线性基函数方法降低一个数量级以上。
  • 适用于复杂行为模式,如状态依赖偏好反转,适合研究多智能体系统的人参考。

针对无限时域平稳平均场博弈(MFG)中的最大因果熵逆强化学习问题,本文在再生核希尔伯特空间(RKHS)中建模未知奖励函数,使能直接从专家示范中推断丰富且可能非线性的奖励结构,突破了传统方法仅限于固定有限基函数线性组合的局限,并避免使用有限时域设定。通过拉格朗日松弛,将原问题转化为无约束对数似然最大化,采用梯度上升算法求解。为证明算法理论一致性,本文证明了相关软贝尔曼算子关于RKHS参数的弗雷歇可微性,从而确保目标函数光滑。在展现实际优势的平均场交通路由游戏中,基于核的方法相较具有相当参数量的线性奖励基线,政策恢复误差降低超过一个数量级。此外,框架扩展至有限时域非平稳情形:发现对数似然重构在该情况下结构上不可行,转而基于丹斯金定理在凸对偶上设计替代梯度下降算法,并建立了光滑性与收敛性保证。

原文摘要 · Abstract (English)

We consider the maximum causal entropy inverse reinforcement learning (IRL) problem for infinite-horizon stationary mean-field games (MFG), in which we model the unknown reward function within a reproducing kernel Hilbert space (RKHS). This allows the inference of rich and potentially nonlinear reward structures directly from expert demonstrations, in contrast to most existing approaches for MFGs that typically restrict the reward to a linear combination of a fixed finite set of basis functions and rely on finite-horizon formulations. We introduce a Lagrangian relaxation that enables us to reformulate the problem as an unconstrained log-likelihood maximization and obtain a solution via a gradient ascent algorithm. To establish the theoretical consistency of the algorithm, we prove the smoothness of the log-likelihood objective through the Fréchet differentiability of the related soft Bellman operators with respect to the parameters in the RKHS. To illustrate the practical advantages of the RKHS formulation, we validate our framework on a mean-field traffic routing game exhibiting state-dependent preference reversal, where the kernel-based method reduces policy recovery error by over an order of magnitude compared to a linear reward baseline with a comparable parameter count. Furthermore, we extend the framework to the finite-horizon non-stationary setting. We demonstrate that the log-likelihood reformulation is structurally unavailable in this regime and instead develop an alternative gradient descent algorithm on the convex dual via Danskin's theorem, establishing smoothness and convergence guarantees.

逆强化学习平均场博弈核方法多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。