arXiv:2601.12707cs.LGstat.ML2026-01

从对手行为反推奖励函数,让博弈决策更可解释。

Decoding Rewards in Competitive Games: Inverse Game Theory with Entropy Regularization

  • 基于熵正则化建立统一框架,从策略中逆向推导奖励。
  • 理论证明奖励函数在特定条件下可识别,算法具样本高效性。
  • 适用于静态与动态博弈,适合研究竞争环境下的决策机制。

估计驱动智能体行为的未知奖励函数是逆强化学习与博弈论中的核心问题。针对双人零和矩阵博弈与马尔可夫博弈,我们提出一种带熵正则化的统一框架,旨在从观测到的玩家策略与动作中重构底层奖励函数。该任务因逆问题固有的模糊性、可行奖励的非唯一性以及观测数据覆盖有限而极具挑战。通过在线性假设下利用量化响应均衡(QRE),我们建立了奖励函数的可识别性。在此理论基础上,提出一种新算法,可从观测动作中学习奖励函数,适用于静态与动态场景,并能融合最大似然估计(MLE)等方法。我们为算法提供了强理论保证,证明其可靠性和样本效率。大量数值实验验证了该框架的实际有效性,为竞争环境中的决策机制提供了新见解。

原文摘要 · Abstract (English)

Estimating the unknown reward functions driving agents' behaviors is of central interest in inverse reinforcement learning and game theory. To tackle this problem, we develop a unified framework for reward function recovery in two-player zero-sum matrix games and Markov games with entropy regularization, where we aim to reconstruct the underlying reward functions given observed players' strategies and actions. This task is challenging due to the inherent ambiguity of inverse problems, the non-uniqueness of feasible rewards, and limited observational data coverage. To address these challenges, we establish the reward function's identifiability using the quantal response equilibrium (QRE) under linear assumptions. Building upon this theoretical foundation, we propose a novel algorithm to learn reward functions from observed actions. Our algorithm works in both static and dynamic settings and is adaptable to incorporate different methods, such as Maximum Likelihood Estimation (MLE). We provide strong theoretical guarantees for the reliability and sample efficiency of our algorithm. Further, we conduct extensive numerical studies to demonstrate the practical effectiveness of the proposed framework, offering new insights into decision-making in competitive environments.

逆强化学习博弈论奖励恢复熵正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。