从无标签混合质量示范中学习多智能体协作,提升数据利用效率。
MisoDICE: Multi-Agent Imitation from Unlabeled Mixed-Quality Demonstrations
- 用大模型与偏好强化学习分步标注示范轨迹质量
- 新算法MisoDICE在专家数据少时仍表现更优
- 适合缺乏高质量示范的多智能体协作场景
我们研究合作式多智能体场景下的离线模仿学习(IL),其中示范数据为未标注的混合质量轨迹,包含专家与次优轨迹。提出两阶段方案:第一阶段结合大语言模型与基于偏好的强化学习,构建渐进式标注流程,区分专家级轨迹;第二阶段引入MisoDICE,一种新型多智能体模仿学习算法,利用标注结果学习鲁棒策略,并解决大规模联合状态-动作空间带来的计算复杂性问题。通过扩展单智能体经典框架DICE,在多智能体设置中引入新的值分解与混合架构,实现凸优化目标,确保全局与局部策略的一致性。在多个标准多智能体强化学习基准上评估,结果表明该方法在专家数据稀缺时表现尤为突出。
原文摘要 · Abstract (English)
We study offline imitation learning (IL) in cooperative multi-agent settings, where demonstrations have unlabeled mixed quality - containing both expert and suboptimal trajectories. Our proposed solution is structured in two stages: trajectory labeling and multi-agent imitation learning, designed jointly to enable effective learning from heterogeneous, unlabeled data. In the first stage, we combine advances in large language models and preference-based reinforcement learning to construct a progressive labeling pipeline that distinguishes expert-quality trajectories. In the second stage, we introduce MisoDICE, a novel multi-agent IL algorithm that leverages these labels to learn robust policies while addressing the computational complexity of large joint state-action spaces. By extending the popular single-agent DICE framework to multi-agent settings with a new value decomposition and mixing architecture, our method yields a convex policy optimization objective and ensures consistency between global and local policies. We evaluate MisoDICE on multiple standard multi-agent RL benchmarks and demonstrate superior performance, especially when expert data is scarce.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。