模型组合的提升有上限,关键在共同失败率。
When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models

- 提出共同失败率β作为组合模型准确率的理论上限。
- 实测67个模型在数学题上共同错误率达5.2%,远超模型预测的2.3%。
- 组合效果受限于不同模型错题不重叠,而非单纯增加模型数量。
多模型语言系统如路由、投票、级联和混合代理被用于超越单模型精度。我们发现其增益受一个领域极少报告的量限制:对于任何输出单一模型答案的策略,准确率无法超过1减去β,其中β是所有模型在同一问题上同时出错的比率。相比之下,常用诊断指标平均成对错误相关性ρ无法识别β:具有相同边际分布和成对相关性的错误规律可能具有不同的全错率。对β的Clopper-Pearson置信区间提供了在训练路由前可实现最大增益的有限样本证明。在来自21家机构的67个模型中,经过四分相关校准的单因素模型仍低估了全错尾部:在开放数学题上,观测到的β为0.052,而基于全部67模型高斯耦合模型预测为0.023,低估约2.5倍(90%置信区间1.7至3.4,k=17)。该现象在执行评分代码任务中再次出现,β为0.079。将GPQA-Diamond问题从选择题改为自由作答形式后,全错率上升至0.127,五名评审员一致性κ为0.73至0.92,表明共同失败源于答案格式而非题目领域。在质量匹配条件下,低ρ异构集成优于高ρ Self-MoA,但在可验证任务中,组合模型很少能超越最优单模型,除非存在强查询级路由信号。增益来自模型错题不重叠,而非模型数量增加。
原文摘要 · Abstract (English)
Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-model accuracy. We show that their gain is capped by a quantity the field rarely reports. For any policy whose output is one member model answer, accuracy cannot exceed one minus beta, where beta is the rate at which every model is wrong on the same query. In contrast, the usual diagnostic, average pairwise error correlation rho, cannot identify beta: error laws with identical marginals and pairwise correlations can have different all-wrong rates. A Clopper-Pearson bound on beta gives a finite-sample certificate on the largest gain any router, vote, or cascade could deliver before training a router. Across 67 models from 21 providers, a tetrachoric-calibrated single-factor model still underprices the all-wrong tail: on open-ended mathematics, observed beta is 0.052 versus 0.023 under the full 67-model Gaussian copula, about 2.5 times underpricing, with 90 percent CI 1.7 to 3.4 and k equals 17. The effect recurs on execution-graded code, where beta is 0.079. Re-asking the same GPQA-Diamond questions in free-response rather than multiple-choice form reopens the tail, with beta 0.127 and a five-judge panel with kappa 0.73 to 0.92, locating co-failure in answer format rather than subject. At matched quality, low-rho heterogeneous ensembles beat high-rho Self-MoA, but on checkable tasks in our pool, combining models rarely beats the single best model without a strong query-level routing signal. Gains come from models failing on different questions, not from adding more models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。