arXiv:2605.14879cs.MAcs.GT2026-05

提出新公平性指标RP,精准衡量多智能体协作的轮换规律性。

Temporal Fair Division in Multi-Agent Systems: From Precise Alternation Metrics to Scalable Coordination Proxies

论文配图:Temporal Fair Division in Multi-Agent Systems: From Precise Alternation Metrics to Scalable Coordination Proxies
图 1 · 摘自论文原文
  • 设计高效度量RP,同时评估等待间隔规律性和访问频率均衡性
  • 实验显示独立训练智能体协调性远差于随机策略,但传统公平指标仍为高分
  • RP速度比精确指标快12-25倍,适合大规模系统评估

众多智能计算与自主系统依赖多个独立、常采用学习机制的智能体反复共享有限资源,如机器人共用工作站、无线设备竞争通信机会、分布式AI智能体争夺计算资源。传统公平性度量仅关注资源分配总量均衡,无法区分有序轮换与不规则访问——后者虽累计结果相似,却导致长时间且不可预测的等待。本文提出计算高效的旋转周期性(Rotational Periodicity, RP)指标,评估成功访问间的等待时间规律性及各智能体访问频率的平衡性。在包含2至10个强化学习智能体的重复阈值拥堵博弈中,对比多种详细交替度量与RP。实验表明,独立训练的智能体协调性显著劣于随机策略智能体,而传统公平指标仍报告良好结果。与此同时,RP能准确复现复杂交替度量的排序,且随智能体数量增加,计算速度提升12至25倍。研究证明,评估多智能体学习系统需引入时序敏感的协调度量,而如RP等高效代理指标使大规模系统评估成为可能。

原文摘要 · Abstract (English)

Many intelligent computing and autonomous systems rely on multiple independent, often learning, agents repeatedly sharing a limited resource. Examples include autonomous robots accessing a shared workstation, wireless devices competing for communication opportunities, and distributed AI agents coordinating access to shared computational resources. While conventional fairness measures assess whether resources are shared equally overall, they cannot distinguish orderly turn-taking from irregular access patterns that produce long and unpredictable waiting times despite similar cumulative outcomes. We introduce Rotational Periodicity (RP), a computationally efficient metric that evaluates both the regularity of waiting times between successful accesses and the balance of access frequencies across agents. We evaluate RP alongside a family of more detailed alternation metrics using a repeated threshold-congestion game in which two to ten reinforcement-learning agents compete for exclusive access to a shared resource. Our experiments reveal that independently trained agents often coordinate substantially worse than random-policy agents, even though conventional fairness metrics consistently report highly favourable outcomes. At the same time, RP closely reproduces the rankings of the more computationally expensive alternation metrics while computing twelve to twenty-five times faster as the number of agents increases. These findings show that evaluating multi-agent learning systems requires temporally aware measures of coordination, not only aggregate outcomes, and that efficient proxy metrics such as RP make this type of evaluation practical for larger intelligent computing systems.

多智能体公平性协调评估强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。