通过重分配奖励提升语言多智能体在跨环境中的协作泛化能力
Advancing Language Multi-Agent Learning with Credit Re-Assignment for Interactive Environment Generalization
- 用大模型动态重分配奖励,取代依赖环境的固定奖励
- 70亿参数系统性能媲美甚至超越闭源强模型
- 适合研究多智能体协作与泛化、希望降低对标注数据依赖的学者
基于大语言模型的智能体在交互环境(如移动操作、网页浏览)中取得显著进展,但现有多智能体系统虽性能优异,却因预设角色和缺乏泛化策略,在跨环境迁移时表现不佳。为解决性能与泛化难以兼得的问题,本文提出CollabUIAgents框架,采用新型多智能体信用重分配(CR)策略:通过大模型生成过程奖励,而非依赖环境特定奖励,并结合合成偏好数据进行训练,使无角色约束的智能体政策能够学习可泛化的协作行为。实验表明,该框架同时提升了性能与跨环境泛化能力;其70亿参数系统在多项任务上达到或超过强大闭源模型水平,且优于指导信用分配的大模型本身。研究还揭示了细粒度信用奖励对环境泛化的有效作用,并提供了将已训练大模型融入多智能体系统的实践路径。
原文摘要 · Abstract (English)
LLM-based agents have made significant advancements in interactive environments, such as mobile operations and web browsing, and other domains beyond computer using. Current multi-agent systems universally excel in performance, compared to single agents, but struggle with generalization across environments due to predefined roles and inadequate strategies for generalizing language agents. The challenge of achieving both strong performance and good generalization has hindered the progress of multi-agent systems for interactive environments. To address these issues, we propose CollabUIAgents, a multi-agent reinforcement learning framework with a novel multi-agent credit re-assignment (CR) strategy, assigning process rewards with LLMs rather than environment-specific rewards and learning with synthesized preference data, in order to foster generalizable, collaborative behaviors among the role-free agents' policies. Empirical results show that our framework improves both performance and cross-environment generalizability of multi-agent systems. Moreover, our 7B-parameter system achieves results on par with or exceed strong closed-source models, and the LLM that guides the CR. We also provide insights in using granular CR rewards effectively for environment generalization, and accommodating trained LLMs in multi-agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。