提出图结构评估框架,揭示多智能体协作中的冗余问题
GEMMAS: Graph-based Evaluation Metrics for Multi Agent Systems
- 将智能体交互建模为有向无环图,分析协作过程
- 发现准确率仅差2.1%时,冗余路径差异达80%
- 适合关注协作效率与可解释性的AI系统研究者
基于语言模型的多智能体系统在协同推理任务中表现优异,但现有评估仅关注最终输出正确性,忽略了通信低效和协调不良导致的冗余推理与高计算成本。本文提出GEMMAS,一种基于图的评估框架,将智能体交互建模为有向无环图。为衡量协作质量,提出两个过程级指标:信息多样性得分(IDS)用于度量智能体间消息的语义多样性,无需路径比(UPR)用于量化冗余推理路径。在五个基准上评估GEMMAS,结果显示在GSM8K上,准确率仅相差2.1%的系统,其IDS差异达12.8%,UPR差异高达80%,表明内部协作存在显著差异。这说明仅依赖结果指标不足以评估多智能体性能,过程级诊断对设计更可解释、资源高效的协作AI系统至关重要。
原文摘要 · Abstract (English)
Multi-agent systems built on language models have shown strong performance on collaborative reasoning tasks. However, existing evaluations focus only on the correctness of the final output, overlooking how inefficient communication and poor coordination contribute to redundant reasoning and higher computational costs. We introduce GEMMAS, a graph-based evaluation framework that analyzes the internal collaboration process by modeling agent interactions as a directed acyclic graph. To capture collaboration quality, we propose two process-level metrics: Information Diversity Score (IDS) to measure semantic variation in inter-agent messages, and Unnecessary Path Ratio (UPR) to quantify redundant reasoning paths. We evaluate GEMMAS across five benchmarks and highlight results on GSM8K, where systems with only a 2.1% difference in accuracy differ by 12.8% in IDS and 80% in UPR, revealing substantial variation in internal collaboration. These findings demonstrate that outcome-only metrics are insufficient for evaluating multi-agent performance and highlight the importance of process-level diagnostics in designing more interpretable and resource-efficient collaborative AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。