arXiv:2603.05789cs.MAcs.GT2026-03被引 2

提出六种新指标,揭示多智能体博弈中表面公平背后的轮流失序问题。

The Coordination Gap: Multi-Agent Alternation Metrics for Temporal Fairness in Repeated Games

  • 设计六种交替性度量(ALT),量化资源争夺的轮换规律性
  • 发现学习策略在交替性上比随机策略差34%-92%,轮换效率仅达理想五分之一
  • 揭示无记忆与有记忆两种失败模式,适合研究公平协作机制的研究者

重复多智能体交互需评估不仅收益分配,还需考察其时间组织。传统基于结果的公平度量会掩盖真实轮流情况。本文以蜂蜜罐游戏(HJG)为范式,引入完美轮换(PA)作为基准轮换模式,并提出六种新型交替性(ALT)度量及对应基准方法。通过Q-learning智能体与分析推导的随机策略对比,发现尽管传统公平性指标高(常超0.9),但学习策略在所有ALT度量上均劣于随机策略,其中CALT差34%-74%,EALT最差达92%;在n=10时,等效轮换效率仅约群体五分之一。EALT还揭示两类失败模式:无记忆时缺陷从n=2的-20%增至n=10的-92%;有记忆时在n=5降至-76%后部分恢复至n=10的-10%。传统效率与公平度量无法识别这些差异。该框架补充了时间公平分配与选取序列文献,能诊断去中心化系统中是否自发产生时间公平性。

原文摘要 · Abstract (English)

Repeated multi-agent interactions require evaluation metrics that capture not only payoff distributions but also their temporal organization. Conventional outcome-based fairness measures can assign similar aggregate scores to temporally distinct coordination patterns, obscuring whether access to a shared resource is genuinely rotating or persistently monopolized. We study this problem in the Honey-Jar Game (HJG), a minimally dynamic repeated threshold-congestion Markov game in which n agents compete for exclusive access to a single high-reward resource. We introduce Perfect Alternation (PA), a reference turn-taking regime corresponding to the n-periodic round-robin picking sequence, together with six novel Alternation (ALT) metrics and a benchmarking methodology mapping ALT values to interpretable PA-equivalent performance. Using Q-learning agents as a minimal adaptive baseline against analytically derived random-policy baselines, we uncover a clear measurement failure: despite deceptively high traditional metrics (e.g., reward fairness often exceeding 0.9), learned policies perform worse than random on every ALT metric, by 34-74% on CALT and up to 92% on EALT, with PA-equivalent coordination falling to roughly one-fifth of the population at n=10. EALT further reveals two distinct failure patterns: without episodic memory (Type-A), the deficit grows from -20% at n=2 to -92% at n=10; with episodic memory (Type-B), it reaches a trough of -76% at n=5 before partially recovering to -10% at n=10. Conventional efficiency and fairness metrics do not reveal these differences. The ALT framework complements the temporal fair division and picking-sequence literature by diagnosing whether temporal fairness emerges spontaneously in decentralized adaptive systems rather than how to enforce it.

多智能体公平性轮换机制博弈论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。