解决动态环境中的多方策略博弈问题,兼顾上下文与对手行为。
Policy Optimization and Statistical Inference for Online Contextual Matrix Games

- 将上下文信息融入多方在线博弈,统一建模动态环境与策略互动。
- 算法实现子线性后悔率,估计的纳什均衡收敛且参数估计渐近正态。
- 适用于竞争定价、广告投放等需同时考虑环境变化与对手反应的场景。
在线决策常面临动态上下文与策略互动的双重挑战。例如,在酒店定价中,需同时考虑市场动态和竞争对手的策略反应。现有方法仅覆盖部分问题:上下文老虎机可利用可观测特征优化单智能体决策,但忽略多方博弈;在线矩阵博弈虽能刻画策略行为,却假设收益固定,未引入上下文信息。为此,本文提出「在线上下文矩阵博弈」框架,并设计「OnGameLearn」算法,高效平衡动作与上下文间的探索与利用。该方法具备统计保障:收益矩阵估计的尾部界、纳什均衡估计的收敛性、参数估计的渐近正态性及子线性后悔率。我们还定义了矩阵博弈中的「策略价值」,并提出一种双重稳健、√T-一致的估计器。在模拟实验与真实酒店定价应用中,OnGameLearn有效应对了策略与上下文交织的复杂决策挑战。
原文摘要 · Abstract (English)
Online decision making often requires navigating a landscape shaped by both dynamic contexts and strategic interactions. In competitive pricing, for example, hotels must account for both dynamic contextual factors and rivals' strategic responses. Existing approaches address only part of this challenge: contextual bandits optimize single-agent decisions using observable features but ignore multi-player interactions, while online matrix games capture strategic behavior through Nash equilibrium but assume fixed payoffs, ignoring contextual information. How should agents act then when strategic payoffs evolve with contextual signals? We introduce \emph{online contextual matrix games} to integrate contextual information into multi-player online games. We further propose \emph{OnGameLearn}, an online learning algorithm that efficiently balances exploration and exploitation across both player actions and contexts. This approach comes with statistical guarantees: tail bounds for the estimated payoff matrix, the convergence of the estimated Nash equilibrium, the asymptotic normality of the parameter estimators, and the sublinear regret bound. We also develop the notion of \emph{policy value} in matrix games and develop a doubly robust, $\sqrt{T}$-consistent estimator for it. Across simulated studies and a real-world hotel pricing application, we find that OnGameLearn effectively navigates the intertwined challenges of strategic and contextual decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。