arXiv:2509.03845cs.LGcs.AI2025-09AAAI被引 9

通过概率上下文变量实现异构智能体的奖励函数逆向学习

Meta-Inverse Reinforcement Learning for Mean Field Games via Probabilistic Context Variables

  • 引入隐变量建模异构任务,无需事先知道上下文信息
  • 在仿真与真实网约车定价场景中优于现有最先进方法
  • 适合处理目标多样但结构相似的多智能体系统

在实际应用中,为大量交互式智能体设计合适的奖励函数极具挑战性。均值场博弈(MFG)中的逆强化学习(IRL)提供了一个从专家演示中推断奖励函数的实用框架。然而,现有方法假设智能体同质化,难以处理现实中普遍存在的异构且未知的目标演示。为此,我们提出一种深度隐变量均值场博弈模型及其配套的IRL方法。关键在于,该方法可在不预先知晓底层上下文或修改原MFG模型的情况下,从不同但结构相似的任务中推断奖励函数。实验在模拟场景和真实世界的时空网约车定价问题上进行,结果表明该方法在均值场博弈的IRL任务中显著优于现有最先进方法。

原文摘要 · Abstract (English)

Designing suitable reward functions for numerous interacting intelligent agents is challenging in real-world applications. Inverse reinforcement learning (IRL) in mean field games (MFGs) offers a practical framework to infer reward functions from expert demonstrations. While promising, the assumption of agent homogeneity limits the capability of existing methods to handle demonstrations with heterogeneous and unknown objectives, which are common in practice. To this end, we propose a deep latent variable MFG model and an associated IRL method. Critically, our method can infer rewards from different yet structurally similar tasks without prior knowledge about underlying contexts or modifying the MFG model itself. Our experiments, conducted on simulated scenarios and a real-world spatial taxi-ride pricing problem, demonstrate the superiority of our approach over state-of-the-art IRL methods in MFGs.

逆强化学习均值场博弈隐变量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。