arXiv:2507.21105cs.IRcs.AI2025-07EMNLP被引 18

用A2A和MCP协议构建多智能体对话系统,实现跨模态信息检索与分析。

AgentMaster: A Multi-Agent Conversational Framework Using A2A and MCP Protocols for Multimodal Information Retrieval and Analysis

  • 融合A2A与MCP双协议,支持智能体间动态协作与灵活通信。
  • 在自然语言交互下完成多任务,BERTScore F1达96.3%,G-Eval为87.1%。
  • 适合需要多工具协同的领域应用,如医疗、金融等复杂场景。

大型语言模型(LLM)与多智能体系统(MAS)的结合显著提升了复杂任务处理能力,但现有系统仍面临智能体间通信、协调及异构工具集成难题。本文提出AgentMaster,一个自研A2A与MCP协议的模块化多协议框架,支持动态协调、灵活通信与快速迭代。通过统一对话接口,系统可自然语言响应多模态查询,完成信息检索、问答与图像分析等任务。人类评估与量化指标验证表明,系统在任务分解、分配、路由及领域相关响应方面表现优异,BERTScore F1达96.3%,LLM-as-a-Judge G-Eval为87.1%。该框架推动了基于MAS的领域专用、协作式、可扩展对话AI的发展。

原文摘要 · Abstract (English)

The rise of Multi-Agent Systems (MAS) in Artificial Intelligence (AI), especially integrated with Large Language Models (LLMs), has greatly facilitated the resolution of complex tasks. However, current systems are still facing challenges of inter-agent communication, coordination, and interaction with heterogeneous tools and resources. Most recently, the Model Context Protocol (MCP) by Anthropic and Agent-to-Agent (A2A) communication protocol by Google have been introduced, and to the best of our knowledge, very few applications exist where both protocols are employed within a single MAS framework. We present a pilot study of AgentMaster, a novel modular multi-protocol MAS framework with self-implemented A2A and MCP, enabling dynamic coordination, flexible communication, and rapid development with faster iteration. Through a unified conversational interface, the system supports natural language interaction without prior technical expertise and responds to multimodal queries for tasks including information retrieval, question answering, and image analysis. The experiments are validated through both human evaluation and quantitative metrics, including BERTScore F1 (96.3%) and LLM-as-a-Judge G-Eval (87.1%). These results demonstrate robust automated inter-agent coordination, query decomposition, task allocation, dynamic routing, and domain-specific relevant responses. Overall, our proposed framework contributes to the potential capabilities of domain-specific, cooperative, and scalable conversational AI powered by MAS.

多智能体对话系统多模态A2A

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。