首次联合解码博弈中奖励与理性参数,突破传统方法对理性度假设的限制。
Blind Inverse Game Theory: Jointly Decoding Rewards and Rationality in Entropy-Regularized Competitive Games
- 提出盲逆博弈框架,同时估计奖励参数与理性温度参数
- 在未知理性度时实现参数唯一识别,收敛速度达最优 $\mathcal{O}(N^{-1/2})$
- 适用于马尔可夫博弈,即使转移动态未知仍表现稳健
基于熵正则化量化响应均衡(QRE)的逆博弈理论(IGT)方法在竞争场景中具有可处理性,但关键假设是代理人的理性参数(温度 $τ$)已知。当 $τ$ 未知时,会引发 $τ$ 与奖励参数 $θ$ 的尺度歧义,导致二者在统计上不可识别。本文提出 Blind-IGT,首个可从观测行为中联合恢复 $θ$ 与 $τ$ 的统计框架。通过分析该双线性反问题,引入归一化约束以消除尺度歧义,建立唯一识别的充要条件。提出高效的归一化最小二乘(NLS)估计器,并证明其达到最优 $\mathcal{O}(N^{-1/2})$ 收敛率。当强可识别性不成立时,通过置信集构造提供部分识别保证。框架扩展至马尔可夫博弈,在转移动态未知情况下仍实现最优收敛率并展现良好实证性能。
原文摘要 · Abstract (English)
Inverse Game Theory (IGT) methods based on the entropy-regularized Quantal Response Equilibrium (QRE) offer a tractable approach for competitive settings, but critically assume the agents' rationality parameter (temperature $τ$) is known a priori. When $τ$ is unknown, a fundamental scale ambiguity emerges that couples $τ$ with the reward parameters ($θ$), making them statistically unidentifiable. We introduce Blind-IGT, the first statistical framework to jointly recover both $θ$ and $τ$ from observed behavior. We analyze this bilinear inverse problem and establish necessary and sufficient conditions for unique identification by introducing a normalization constraint that resolves the scale ambiguity. We propose an efficient Normalized Least Squares (NLS) estimator and prove it achieves the optimal $\mathcal{O}(N^{-1/2})$ convergence rate for joint parameter recovery. When strong identifiability conditions fail, we provide partial identification guarantees through confidence set construction. We extend our framework to Markov games and demonstrate optimal convergence rates with strong empirical performance even when transition dynamics are unknown.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。