研究团队中多样性能带来优势的条件,提出可验证的奖励设计原则。
When Is Diversity Rewarded in Cooperative Multi-Agent Learning?
- 用双层聚合算子建模任务分配,通过凸性判断是否奖励多样性。
- 在多智能体强化学习中发现,特定奖励结构能显著提升异质团队表现。
- 提出新算法HetGPS,自动寻找有利于多样性的环境参数配置。
团队在机器人、自然和人类社会中的成功常依赖于不同专长成员的分工协作;然而,为何某些情况下多样性优于同质团队仍缺乏系统解释。本文聚焦多智能体任务分配问题,从奖励设计角度探讨:何种目标函数最适合异质团队?首先,在非空间即时设定下,构建两层聚合机制——内层将N个智能体在各任务上的努力分配映射为任务得分,外层将M个任务得分合并为全局团队奖励。理论证明,该机制的曲率决定了多样性能否提升收益,对广泛奖励族而言,这简化为一个凸性检验。其次,针对具身化、时间延展的智能体需学习努力分配策略的情形,采用多智能体强化学习(MARL)作为计算框架,提出梯度驱动的异质性增益参数搜索(HetGPS)算法,优化未充分定义的MARL环境参数空间,以发现多样性占优的情境。在多个环境中,HetGPS重现了理论预测的最优奖励区间,既验证了算法有效性,也连接了理论洞察与实际奖励设计。结果揭示了行为多样性产生可观收益的条件。
原文摘要 · Abstract (English)
The success of teams in robotics, nature, and society often depends on the division of labor among diverse specialists; however, a principled explanation for when such diversity surpasses a homogeneous team is still missing. Focusing on multi-agent task allocation problems, we study this question from the perspective of reward design: what kinds of objectives are best suited for heterogeneous teams? We first consider an instantaneous, non-spatial setting where the global reward is built by two generalized aggregation operators: an inner operator that maps the $N$ agents' effort allocations on individual tasks to a task score, and an outer operator that merges the $M$ task scores into the global team reward. We prove that the curvature of these operators determines whether heterogeneity can increase reward, and that for broad reward families this collapses to a simple convexity test. Next, we ask what incentivizes heterogeneity to emerge when embodied, time-extended agents must learn an effort allocation policy. To study heterogeneity in such settings, we use multi-agent reinforcement learning (MARL) as our computational paradigm, and introduce Heterogeneity Gain Parameter Search (HetGPS), a gradient-based algorithm that optimizes the parameter space of underspecified MARL environments to find scenarios where heterogeneity is advantageous. Across different environments, we show that HetGPS rediscovers the reward regimes predicted by our theory to maximize the advantage of heterogeneity, both validating HetGPS and connecting our theoretical insights to reward design in MARL. Together, these results help us understand when behavioral diversity delivers a measurable benefit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。