arXiv:2604.22760cs.IRcs.AI2026-04

量化大模型调用API时的差异,发现抽象任务下模型意见分歧严重。

Quantifying Divergence in Inter-LLM Communication Through API Retrieval and Ranking

  • 构建统一框架,通过集合、排序和共识指标衡量模型间差异。
  • 平均重叠度约0.50,肯德尔等级相关系数约0.45,抽象任务分歧更大。
  • 适合多智能体系统设计者,用于识别潜在协同风险与安全漏洞。

大型语言模型(LLMs)越来越多地作为自主代理,通过外部API执行复杂任务。然而其可靠性与一致性尚未被充分理解。本文提出一个统一基准框架,量化不同模型在相同任务下进行API发现与排序时的差异。在15个典型API领域和5个主流模型家族中,采用基于集合、排序和共识的多种指标(如平均重叠度、杰卡德相似度、秩偏重叠度、肯德尔tau、肯德尔W、克朗巴赫α)评估成对及群体级一致性。结果表明整体对齐度中等(平均重叠度约0.50,肯德尔tau约0.45),但存在显著领域依赖:结构化任务(如天气查询、语音转文本)较稳定,而开放式任务(如情感分析)分歧显著更高。波动性与共识分析显示,一致性集中于数据驱动型领域,在抽象推理任务中明显下降。这些发现支持多智能体系统中的可靠性感知编排,共识加权可提升异构模型间的协作能力。除性能评估外,研究揭示了多智能体协同中的系统性失效模式:表面一致可能掩盖关键动作排序的不稳定性。这种隐藏分歧构成部署前的安全隐患,亟需诊断性基准以实现早期检测。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly operate as autonomous agents that reason over external APIs to perform complex tasks. However, their reliability and agreement remain poorly characterized. We present a unified benchmarking framework to quantify inter-LLM divergence, defined as the extent to which models differ in API discovery and ranking under identical tasks. Across 15 canonical API domains and 5 major model families, we measure pairwise and group-level agreement using set-, rank-, and consensus-based metrics including Average Overlap, Jaccard similarity, Rank-Biased Overlap, Kendall's tau, Kendall's W, and Cronbach's alpha. Results show moderate overall alignment (AO about 0.50, tau about 0.45) but strong domain dependence: structured tasks (Weather, Speech-to-Text) are stable, while open-ended tasks (Sentiment Analysis) exhibit substantially higher divergence. Volatility and consensus analyses reveal that coherence clusters around data-bound domains and degrades for abstract reasoning tasks. These insights enable reliability-aware orchestration in multi-agent systems, where consensus weighting can improve coordination among heterogeneous LLMs. Beyond performance benchmarking, our results reveal systematic failure modes in multi-agent LLM coordination, where apparent agreement can mask instability in action-relevant rankings. This hidden divergence poses a pre-deployment safety risk and motivates diagnostic benchmarks for early detection.

大模型多智能体一致性评估安全风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。