arXiv:2502.07350cs.AI2025-02ICML被引 43

用动态知识感知的强化学习,让多个智能体协作更聪明、更省资源。

KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent Systems

  • 构建三维语义距离模型,理解智能体间知识差异
  • 双适应机制持续优化专家能力,响应环境变化
  • 基于知识的汤普森采样策略,高效选择最优协作方案

随着大语言模型规模扩大带来高昂成本,多智能体系统成为有前景的替代方案,但面临知识静态假设和协作效率低下的挑战。本文提出知识感知贝叶斯博弈(KABB)框架,通过语义理解与动态适应提升多智能体协作效率。该框架包含三项核心创新:用于深层语义理解的三维知识距离模型、支持持续专家优化的双适应机制,以及实现高效专家选择的知识感知汤普森采样策略。大量实验表明,KABB在多智能体协作中实现了成本与性能的最优平衡,在保持高性能的同时显著降低计算开销。

原文摘要 · Abstract (English)

As scaling large language models faces prohibitive costs, multi-agent systems emerge as a promising alternative, though challenged by static knowledge assumptions and coordination inefficiencies. We introduces Knowledge-Aware Bayesian Bandits (KABB), a novel framework that enhances multi-agent system coordination through semantic understanding and dynamic adaptation. The framework features three key innovations: a three-dimensional knowledge distance model for deep semantic understanding, a dual-adaptation mechanism for continuous expert optimization, and a knowledge-aware Thompson Sampling strategy for efficient expert selection. Extensive evaluation demonstrates KABB achieves an optimal cost-performance balance, maintaining high performance while keeping computational demands relatively low in multi-agent coordination.

多智能体贝叶斯优化知识协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。