首次给出一般和博弈中栈桥策略Q值迭代的有限时间收敛保证
Finite-Time Analysis of Q-Value Iteration for General-Sum Stackelberg Games
- 基于控制论视角建模栈桥学习为切换系统,引入松弛策略条件
- 构建上下界比较系统,得到Q函数的有限时间误差上界
- 为多智能体强化学习提供理论新视角,适合关注博弈学习理论的研究者
强化学习在单智能体场景中已取得理论与实证双重成功,但将其扩展到一般和博弈中的多智能体强化学习仍具挑战。本文从控制论角度研究双人一般和马尔可夫博弈中栈桥Q值迭代的收敛性。提出适配栈桥设定的松弛策略条件,并将学习动态建模为切换系统。通过构造上下界比较系统,建立了Q函数的有限时间误差界,并刻画了其收敛性质。结果提供了栈桥学习的新控制论视角。据作者所知,本论文首次在栈桥交互下为一般和马尔可夫博弈中的Q值迭代提供了有限时间收敛保证。
原文摘要 · Abstract (English)
Reinforcement learning has been successful both empirically and theoretically in single-agent settings, but extending these results to multi-agent reinforcement learning in general-sum Markov games remains challenging. This paper studies the convergence of Stackelberg Q-value iteration in two-player general-sum Markov games from a control-theoretic perspective. We introduce a relaxed policy condition tailored to the Stackelberg setting and model the learning dynamics as a switching system. By constructing upper and lower comparison systems, we establish finite-time error bounds for the Q-functions and characterize their convergence properties. Our results provide a novel control-theoretic perspective on Stackelberg learning. Moreover, to the best of the authors' knowledge, this paper offers the first finite-time convergence guarantees for Q-value iteration in general-sum Markov games under Stackelberg interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。