arXiv:2502.08365cs.LGcs.AI2025-02NeurIPS被引 5

提出可扩展的无监督多智能体强化学习方法,实现高效探索。

Towards Principled Unsupervised Multi-Agent Reinforcement Learning

  • 采用混合熵最大化作为无监督目标,引导智能体自主探索
  • 在复杂环境中验证,显著提升下游任务学习效率
  • 适合需要自适应协作的多智能体系统研究者

在强化学习中,无监督预训练旨在不依赖任务奖励的情况下预先训练策略,以支持后续下游任务的高效学习。单智能体场景下该问题已基本明确:通过最大化状态分布熵实现无任务探索。但在多智能体场景中,这一问题仍不清晰。本文首先分析不同形式化方案的优劣,揭示即使理论上可行,实际中也面临挑战。随后提出一种可扩展、去中心化、基于信任区域的策略搜索算法,解决实际应用难题。数值实验验证了理论结论,并表明优化混合熵能在可计算性与性能间取得良好平衡,为复杂环境下基于无任务探索的多智能体强化学习提供可行路径。

原文摘要 · Abstract (English)

In reinforcement learning, we typically refer to unsupervised pre-training when we aim to pre-train a policy without a priori access to the task specification, i.e. rewards, to be later employed for efficient learning of downstream tasks. In single-agent settings, the problem has been extensively studied and mostly understood. A popular approach, called task-agnostic exploration, casts the unsupervised objective as maximizing the entropy of the state distribution induced by the agent's policy, from which principles and methods follow. In contrast, little is known about it in multi-agent settings, which are ubiquitous in the real world. What are the pros and cons of alternative problem formulations in this setting? How hard is the problem in theory, how can we solve it in practice? In this paper, we address these questions by first characterizing those alternative formulations and highlighting how the problem, even when tractable in theory, is non-trivial in practice. Then, we present a scalable, decentralized, trust-region policy search algorithm to address the problem in practical settings. Finally, we provide numerical validations to both corroborate the theoretical findings and pave the way for unsupervised multi-agent reinforcement learning via task-agnostic exploration in challenging domains, showing that optimizing for a specific objective, namely mixture entropy, provides an excellent trade-off between tractability and performances.

多智能体无监督学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。