arXiv:2503.09309cs.LGcs.AI2025-03被引 2

在模型不确定下,用激励机制引导大规模群体行为收敛。

Steering No-Regret Agents in MFGs under Model Uncertainty

  • 基于无后悔学习设计乐观探索算法,动态调整激励策略。
  • 累积行为偏差与目标偏差的差距呈次线性增长,性能逼近理论最优。
  • 适合大规模、未知动态系统的智能体引导,如交通调度或市场调控。

激励设计是一种通过提供超出内在奖励的额外支付来引导智能体学习动态向期望结果演进的框架。然而,现有工作大多局限于有限且小规模的智能体集合,或假设对博弈过程有完全认知,难以适用于包含大规模群体和模型不确定性的现实场景。为此,本文研究在密度无关转移的平均场博弈(MFGs)中设计引导激励的问题,其中转移动态和内在奖励函数均未知。该设定带来非平凡挑战:中介需激励智能体探索以学习模型,同时引导其收敛至期望行为,且激励成本不能过高。假设智能体呈现无(非自适应)后悔行为,本文提出新颖的乐观探索算法。理论上,建立了智能体行为与目标行为间累积差距的次线性后悔保证。在引导成本方面,证明总激励支出仅产生次线性超额成本,可与将目标策略稳定为均衡的基准策略相竞争。本工作为大规模群体系统在不确定性下的行为引导提供了有效框架。

原文摘要 · Abstract (English)

Incentive design is a popular framework for guiding agents' learning dynamics towards desired outcomes by providing additional payments beyond intrinsic rewards. However, most existing works focus on a finite, small set of agents or assume complete knowledge of the game, limiting their applicability to real-world scenarios involving large populations and model uncertainty. To address this gap, we study the design of steering rewards in Mean-Field Games (MFGs) with density-independent transitions, where both the transition dynamics and intrinsic reward functions are unknown. This setting presents non-trivial challenges, as the mediator must incentivize the agents to explore for its model learning under uncertainty, while simultaneously steer them to converge to desired behaviors without incurring excessive incentive payments. Assuming agents exhibit no(-adaptive) regret behaviors, we contribute novel optimistic exploration algorithms. Theoretically, we establish sub-linear regret guarantees for the cumulative gaps between the agents' behaviors and the desired ones. In terms of the steering cost, we demonstrate that our total incentive payments incur only sub-linear excess, competing with a baseline steering strategy that stabilizes the target policy as an equilibrium. Our work presents an effective framework for steering agents behaviors in large-population systems under uncertainty.

平均场博弈激励设计无后悔学习大规模系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。