arXiv:2509.22426cs.LGcs.GT2025-09NeurIPS被引 1

延迟反馈下,新算法让多智能体博弈更快收敛到均衡

Learning from Delayed Feedback in Games via Extra Prediction

  • 用加权预测补偿时间延迟,改进原有乐观策略算法
  • 单步延迟就让原算法性能下降,新算法能恢复常数社会遗憾
  • 适合研究多智能体强化学习与博弈论的学者参考

本文研究了博弈学习中的时延反馈问题。由于多智能体独立优化策略,常出现优化偏差。现有乐观跟随正则化领导者(OFTRL)算法依赖未来奖励预测来缓解该问题,但观测过去奖励存在时延会阻碍预测效果。本文首次证明,即使单步延迟也会导致OFTRL在社会遗憾和收敛性方面性能下降。为此提出加权OFTRL(WOFTRL),将下一时刻奖励的预测向量乘以权重因子 $n$。研究发现,当乐观权重超过时延步数时,可消除时延影响:在一般和博弈中社会遗憾保持有界,在多项式零和博弈中策略最后迭代收敛至纳什均衡。理论分析得到实验验证。

原文摘要 · Abstract (English)

This study raises and addresses the problem of time-delayed feedback in learning in games. Because learning in games assumes that multiple agents independently learn their strategies, a discrepancy in optimization often emerges among the agents. To overcome this discrepancy, the prediction of the future reward is incorporated into algorithms, typically known as Optimistic Follow-the-Regularized-Leader (OFTRL). However, the time delay in observing the past rewards hinders the prediction. Indeed, this study firstly proves that even a single-step delay worsens the performance of OFTRL from the aspects of social regret and convergence. This study proposes the weighted OFTRL (WOFTRL), where the prediction vector of the next reward in OFTRL is weighted $n$ times. We further capture an intuition that the optimistic weight cancels out this time delay. We prove that when the optimistic weight exceeds the time delay, our WOFTRL recovers the good performances that social regret is constant in general-sum normal-form games, and the strategies last-iterate converge to the Nash equilibrium in poly-matrix zero-sum games. The theoretical results are supported and strengthened by our experiments.

博弈学习时延反馈纳什均衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。