arXiv:2506.05577cs.LGcs.AI2025-06AAAI

智能体通过任务相似性共享策略,加速学习并解决复杂任务。

Policy Search, Retrieval, and Composition via Task Similarity in Collaborative Agentic Systems

  • 基于任务嵌入的相似度与性能信号选择最优策略
  • 模块化神经结构支持策略复用与快速集成
  • 可减少任务干扰,适合多智能体协作场景

自主智能体系统旨在设定目标、主动适应变化,并通过持续经验优化行为。面对多重未知任务时,若能共享已学习的策略知识,将显著提升效率。然而,如何从其他智能体中查询、选择并整合策略仍不明确。本文提出MOSAIC算法,通过(1)基于Wasserstein任务嵌入的余弦相似度与性能信号进行知识筛选;(2)利用掩码实现模块化、可迁移的神经表示;(3)策略整合、组合与微调。实验表明,MOSAIC在学习速度和整体表现上优于孤立学习与全局共享方法,在部分任务中甚至解决了孤立智能体无法完成的问题。结果还显示,有选择性的目标驱动复用可降低任务干扰。此外观察到自组织现象:解决简单任务的智能体通过知识共享加速了复杂任务的学习。

原文摘要 · Abstract (English)

Agentic AI aims to create systems that set their own goals, adapt proactively to change, and refine behavior through continuous experience. Recent advances suggest that, when facing multiple and unforeseen tasks, agents could benefit from sharing machine-learned knowledge and reusing policies that have already been fully or partially learned by other agents. However, how to query, select, and retrieve policies from a pool of agents, and how to integrate such policies remains a largely unexplored area. This study explores how an agent decides what knowledge to select, from whom, and when and how to integrate it in its own policy in order to accelerate its own learning. The proposed algorithm, \emph{Modular Sharing and Composition in Collective Learning} (MOSAIC), improves learning in agentic collectives by combining (1) knowledge selection using performance signals and cosine similarity on Wasserstein task embeddings, (2) modular and transferable neural representations via masks, and (3) policy integration, composition and fine-tuning. MOSAIC outperforms isolated learners and global sharing approaches in both learning speed and overall performance, and in some cases solves tasks that isolated agents cannot. The results also demonstrate that selective, goal-driven reuse leads to less susceptibility to task interference. We also observe the emergence of self-organization, where agents solving simpler tasks accelerate the learning of harder ones through shared knowledge.

智能体协作策略复用自主学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。