用流模型将博弈空间映射到球面流形,实现高效在线学习。
Riemannian Manifold Learning for Stackelberg Games with Neural Flow Representations
- 通过神经归一化流学习动作空间到球面流形的微分同胚映射。
- 在流形上利用线性奖励函数,使线性老虎机算法可直接应用。
- 理论证明了后悔率有界,适合多智能体博弈与安全优化场景。
我们提出一种新型在线学习框架,用于处理领导者-跟随者结构的广义和博弈,其中两名智能体按顺序交互。核心是通过神经归一化流学习一个微分同胚,将联合动作空间映射至光滑的球面黎曼流形(称作斯塔克尔伯格流形),该映射生成可处理的等平面子空间,从而支持高效的在线学习。由于代理在斯塔克尔伯格流形上的收益函数呈线性特性,本方法可直接应用线性老虎机算法。我们为流形上的后悔最小化提供了严格的理论基础,并建立了学习斯塔克尔伯格均衡的简单后悔上界。该框架将流形学习引入博弈论,揭示了神经归一化流在多智能体学习中的强大潜力。实验表明,本方法在网络安全与供应链优化等任务中优于标准基线。
原文摘要 · Abstract (English)
We present a novel framework for online learning in Stackelberg general-sum games, where two agents, the leader and follower, engage in sequential turn-based interactions. At the core of this approach is a learned diffeomorphism that maps the joint action space to a smooth spherical Riemannian manifold, referred to as the Stackelberg manifold. This mapping, facilitated by neural normalizing flows, ensures the formation of tractable isoplanar subspaces, enabling efficient techniques for online learning. Leveraging the linearity of the agents' reward functions on the Stackelberg manifold, our construct allows the application of linear bandit algorithms. We then provide a rigorous theoretical basis for regret minimization on the learned manifold and establish bounds on the simple regret for learning Stackelberg equilibrium. This integration of manifold learning into game theory uncovers a previously unrecognized potential for neural normalizing flows as an effective tool for multi-agent learning. We present empirical results demonstrating the effectiveness of our approach compared to standard baselines, with applications spanning domains such as cybersecurity and economic supply chain optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。