让自利智能体学会合作,通过建模彼此学习过程。
Multi-agent cooperation through learning-aware policy gradients

- 用无偏高阶导数策略梯度,建模其他智能体的学习动态。
- 在社交困境中实现高效合作,最高收益超基线30%。
- 适用于需长期协作的复杂任务,适合多智能体系统研究者。
自利个体常无法合作,这是多智能体学习的核心挑战。如何使自利、独立学习的智能体实现协作?近期研究发现,学习感知型智能体若能建模彼此的学习动态,可在特定任务中达成合作。本文提出首个无偏、无需高阶导数的策略梯度算法,用于学习感知强化学习,该算法考虑其他智能体基于多次噪声试验进行试错学习。我们进一步利用高效的序列模型,使行为基于包含其他智能体学习痕迹的长历史观察。使用该算法训练长上下文策略,在标准社交困境任务中实现了合作行为和高回报,包括一个需要时间延展动作协调的挑战性环境。最后,我们从迭代囚徒困境推导出一种新解释,阐明自利学习感知智能体何时以及如何产生合作。
原文摘要 · Abstract (English)
Self-interested individuals often fail to cooperate, posing a fundamental challenge for multi-agent learning. How can we achieve cooperation among self-interested, independent learning agents? Promising recent work has shown that in certain tasks cooperation can be established between learning-aware agents who model the learning dynamics of each other. Here, we present the first unbiased, higher-derivative-free policy gradient algorithm for learning-aware reinforcement learning, which takes into account that other agents are themselves learning through trial and error based on multiple noisy trials. We then leverage efficient sequence models to condition behavior on long observation histories that contain traces of the learning dynamics of other agents. Training long-context policies with our algorithm leads to cooperative behavior and high returns on standard social dilemmas, including a challenging environment where temporally-extended action coordination is required. Finally, we derive from the iterated prisoner's dilemma a novel explanation for how and when cooperation arises among self-interested learning-aware agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。