发现多智能体协作中选择环节是关键瓶颈,决定多样性是否提升效果。
When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines
- 提出选择瓶颈模型,揭示多样性何时有益或有害
- 实验证明选优策略胜过合成聚合,胜率高达81%
- 适合关注智能体协作设计与评估的研究者
多智能体大模型流水线在团队多样性是否提升输出质量上存在矛盾:异构混合智能体团队优于单模型,但同构自混合团队在基于合成的聚合下始终占优。本文提出通过识别‘选择瓶颈’来解决这一矛盾——即聚合质量存在一个临界阈值,决定多样性是增益还是损害。我们推导出该阈值 $s^*$ 的闭式表达(命题1),并设计了跨42个任务、7个类别的实验($N=210$)。采用基于裁判的选择机制,多样团队胜率为0.810,远超单模型基线(0.512);而合成方式在42项任务中无一胜于基线。裁判评价显示,选优策略比合成方式胜率高出0.631。独立评委验证结果方向一致(斯皮尔曼相关ρ=0.90)。初步证据表明引入较弱模型可提升性能并降低成本($p < 10^{-4}$,未预注册)。结果表明,在单轮生成-选择流程中,选择器质量可能比生成器多样性更具影响力。
原文摘要 · Abstract (English)
Multi-agent LLM pipelines produce contradictory evidence on whether team diversity improves output quality: heterogeneous Mixture-of-Agents teams outperform single models, yet homogeneous Self-MoA teams consistently win under synthesis-based aggregation. We propose a resolution by identifying the selection bottleneck -- a crossover threshold in aggregation quality that determines whether diversity helps or hurts. Under this model, we obtain a closed-form crossover threshold $s^*$ (Proposition 1) that separates the regimes where diversity helps and hurts. In a targeted experiment spanning 42 tasks across 7 categories ($N=210$), a diverse team with judge-based selection achieves a win rate of 0.810 against a single-model baseline, while a homogeneous team scores 0.512 -- near chance (Glass's $Δ= 2.07$). Judge-based selection outperforms MoA-style synthesis by $Δ_{\mathrm{WR}} = +0.631$ -- the synthesis approach is preferred over the baseline in zero of 42 tasks by the judge panel. A decoupled evaluation with independent judges confirms all directional findings (Spearman $ρ= 0.90$). Exploratory evidence suggests that including a weaker model improves performance while reducing cost ($p < 10^{-4}$, not pre-registered). Our results suggest that selector quality may be a more impactful design lever than generator diversity in single-round generate-then-select pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。