arXiv:2510.01499cs.LGcs.AI2025-10中稿 · ICML被引 21

用高阶信息优化大模型答案聚合,提升决策可靠性

Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information

  • 基于一阶与二阶信息设计新型聚合算法
  • 在多个数据集上均超越传统多数投票方法
  • 无需训练,适合多模型协作场景

随着多智能体大语言模型推理的快速发展,如何有效聚合多个大模型的答案已成为核心挑战。标准多数投票将所有答案等同看待,未能考虑模型间的潜在异质性和相关性。本文提出两种新聚合算法:最优权重(OW)和逆意外流行度(ISP),利用一阶与二阶信息。理论分析表明,在温和假设下,这些方法可有效缓解多数投票的固有缺陷,提升集体决策的可靠性。我们在合成数据集、UltraFeedback 和 MMLU 等主流微调基准,以及真实医疗场景 ARMMAN 上进行实证验证。结果表明,所提算法在各场景中均持续优于标准基线,构建了一个无需训练的稳健多智能体大模型聚合框架。

原文摘要 · Abstract (English)

With the rapid progress of multi-agent large language model (LLM) reasoning, how to effectively aggregate answers from multiple LLMs has emerged as a fundamental challenge. Standard majority voting treats all answers equally, failing to consider latent heterogeneity and correlation across models. In this work, we design two new aggregation algorithms called Optimal Weight (OW) and Inverse Surprising Popularity (ISP), leveraging both first-order and second-order information. Our theoretical analysis shows these methods provably mitigate inherent limitations of majority voting under mild assumptions, leading to more reliable collective decisions. We empirically validate our algorithms on synthetic datasets, popular LLM fine-tuning benchmarks such as UltraFeedback and MMLU, and a real-world healthcare setting ARMMAN. Our algorithms consistently outperform standard baselines, establishing a robust, training-free framework for effective multi-agent LLM aggregation.

大模型聚合多智能体推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。