arXiv:2601.03451stat.MLcs.AI2026-01

为多智能体学习设计激励机制,实现渐近最优社会福利。

Microeconomic Foundations of Multi-Agent Learning

  • 分两阶段设计激励机制,先估转移量再引导长期学习
  • 在弱理性与探索条件下实现亚线性社会福利遗憾
  • 适合关注市场中AI安全与福利对齐的研究者

现代AI系统越来越多地运行在数据、行为和激励内生的市场与制度中。本文通过研究带有战略外部性的马尔可夫决策过程中的委托-代理互动,为多智能体学习建立经济基础,其中委托方与代理方均随时间学习。提出一种两阶段激励机制:首先估计可实施的转移量,再利用其引导长期动态;在温和的基于后悔的理性与探索条件下,该机制实现了亚线性社会福利遗憾,从而达到渐近最优福利。模拟显示,即使粗略的激励也能纠正状态依赖外部性下的低效学习,凸显了在市场与保险等场景中进行激励感知设计对安全与福利对齐AI的必要性。

原文摘要 · Abstract (English)

Modern AI systems increasingly operate inside markets and institutions where data, behavior, and incentives are endogenous. This paper develops an economic foundation for multi-agent learning by studying a principal-agent interaction in a Markov decision process with strategic externalities, where both the principal and the agent learn over time. We propose a two-phase incentive mechanism that first estimates implementable transfers and then uses them to steer long-run dynamics; under mild regret-based rationality and exploration conditions, the mechanism achieves sublinear social-welfare regret and thus asymptotically optimal welfare. Simulations illustrate how even coarse incentives can correct inefficient learning under stateful externalities, highlighting the necessity of incentive-aware design for safe and welfare-aligned AI in markets and insurance.

多智能体激励机制学习理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。