让多个大模型接力补短板,效率比单个模型高得多。
Harnessing the Wisdom of LLM Crowds through Complementarity-Driven Iterative Collaboration

- 模型按需接力,后一个专攻前一个的薄弱环节。
- 在多个测试中表现优于单模型和集成方法。
- 适合需要低成本、自部署的智能系统设计。
大型语言模型(LLMs)在企业应用中日益普及,但单个模型的能力受限于自身特性。这些异质性局限带来了挑战,也创造了机会:通过策略性协调多个模型,可实现超越单一模型的集体智能。现有方法预先固定模型组合方式,忽略了复杂问题求解中互补性的动态变化。受群体智慧启发,本文将多模型智能重新定义为接力式互补:每个后续模型根据前序输出识别出的具体瓶颈进行选择。为此提出WILC(Wisdom Integration of LLM Crowds)框架,基于两个设计原则:一是迭代反思与优化,保持状态延续性;二是基于双门控机制的互补驱动选型——前瞻性互补契合度(PCF)识别最适胜任者,事后悔补增益(PCG)评估转换是否提升解决方案。在四个不同基准测试中,WILC均优于现有方法,包括单模型自我修正、集成方法及查询路由。在标准定价下,其性能接近GPT-5.2平均水平,但每请求成本仅为约1/7,且支持自托管部署,保障数据主权。本研究将群体智慧理论从静态聚合扩展至序列化人工智能互补,并提供可迁移的多智能体协同设计原则。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in enterprise settings, yet individual models remain bounded by model-specific capability limitations. These heterogeneous boundaries pose a deployment challenge, but also create an opportunity: strategically coordinating multiple LLMs may unlock collective intelligence exceeding any single model. Existing approaches fix how models are combined in advance, overlooking the dynamic, state-dependent role of complementarity in complex problem solving. Drawing on the wisdom-of-crowds paradigm, we reconceptualize collective LLM intelligence as relay-style complementarity: a sequential process in which each successor model is selected to address the specific bottleneck identified in its predecessor's output. To operationalize this, we propose WILC (Wisdom Integration of LLM Crowds), a framework grounded in two design principles. First, iterative reflection-and-refinement establishes a state-preserving workflow through which models diagnose and refine prior outputs. Second, complementarity-driven model selection governs transitions via a dual-gate mechanism: prospective complementarity fit (PCF) identifies the worker most suited to the current bottleneck, while posterior complementarity gain (PCG) evaluates whether the selected transition improves the evolving solution. Experiments across four diverse benchmarks show that WILC outperforms existing approaches, including single-model self-refinement, ensemble methods, and query-routing methods. Under standardized pricing assumptions, WILC matches the average benchmark performance of GPT-5.2 at roughly 7 times lower estimated per-query cost, while facilitating data sovereignty through self-hosted deployment. This study extends wisdom-of-crowds theory from static aggregation to sequential AI complementarity and provides transferable design principles for multi-AI coordination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。