arXiv:2605.11042cs.GTcs.AI2026-05

无需模型知识,让新用户在公平资源系统中自主学习最优策略。

Towards Model-Free Learning in Dynamic Population Games: An Application to Karma Economies

论文配图:Towards Model-Free Learning in Dynamic Population Games: An Application to Karma Economies
图 1 · 摘自论文原文
  • 用深度Q网络让新加入者在不掌握规则的情况下学习策略。
  • 理论证明策略偏离最优解的误差随群体规模增大而减小。
  • 首次实现无模型下动态群体博弈的均衡收敛,适合分布式系统设计者。

动态群体博弈(DPG)为大规模自利个体间的策略互动提供了可计算的框架,已被成功应用于设计“信誉经济”——一种公平的非货币资源分配机制。尽管理论优势明显,现有工具依赖完整游戏模型且需中心化计算,难以适用于代理仅能访问自身经验的真实场景。本文首次研究信誉型DPG中的无模型均衡学习:首先分析新代理在已有稳定纳什均衡(SNE)环境中,通过深度Q网络(DQN)无模型学习的情形;基于近期DQN收敛结果,建立子最优性界,包含约 $O(1/ ext{√}N_s)$ 的近似误差和约 $O(1/N)$ 的均场扰动误差,其中 $N_s$ 为经验回放大小,$N$ 为种群规模。其次,针对从零开始学习SNE的挑战,实验证明结合深度强化学习、虚构博弈与平滑策略迭代,可实现模型无关的收敛至接近中心计算的SNE。这些成果支持将信誉经济作为实际公平资源分配工具的愿景。

原文摘要 · Abstract (English)

Dynamic Population Games (DPGs) provide a tractable framework for modeling strategic interactions in large populations of self-interested agents, and have been successfully applied to the design of Karma economies, a class of fair non-monetary resource allocation mechanisms. Despite their appealing theoretical properties, existing computational tools for DPGs assume full knowledge of the game model and operate in a centralized fashion, limiting their applicability in realistic settings where agents have access only to their own private experience. This paper takes a step towards addressing this gap by studying model-free equilibrium learning in Karma DPGs. First, we analyze the setting in which a novel agent joins a Karma DPG already at its Stationary Nash Equilibrium (SNE) and learns a policy via Deep Q-Networks (DQN) without knowledge of the game model. Leveraging recent convergence results for DQN, we establish a suboptimality bound consisting of a DQN approximation error of order $O(1/\sqrt{N_s})$ and a mean field perturbation error of order $O(1/N)$, where $N_s$ is the replay buffer size and $N$ is the population size. Second, we consider the challenging problem of learning the SNE from scratch. We show empirically that combining deep RL with fictitious play and smoothed policy iteration allows agents to converge, in a model-free fashion, to a configuration close to the centrally computed SNE. Together, these contributions support the vision of Karma economies as practical tools for fair resource allocation.

无模型学习信誉经济强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。