检验大模型集成中多样性度量是否真测多样,发现多数度量实为能力反映。
Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles

- 在30个大模型组合中控制能力,验证多样性度量的可靠性。
- 仅9.98%的三模型组合胜过最强单体,说明多数集成无实际增益。
- 多数多样性指标与模型能力高度相关,不适合独立评估多样性。
主流观点认为大模型集成通过多样性提升性能,而多样性度量被用于筛选组合模型。本文在MMLU-Pro(31,900个子集,30个LLM)和TruthfulQA(29个)上,以明确的能力控制条件,审计五种多样性度量对集成性能增益(即多数投票优于最强模型)的预测能力。结果表明:第一,潜在互补性普遍存在——所有子集均存在理论上可提升的“皇牌增益”,但简单投票仅在9.98%的三模型标准子集中超越最强成员(若允许保留最佳模型,则升至18.71%),规模为2-4的总体增益率仅为1.27%,部分源于偶数规模投票的确定性行为;第二,严格多样性(联合正确性代理)与1减去平均准确率高度共线(大小为3时斯皮尔曼等级相关系数达+0.991/+0.988),原始多样性关联严重受能力影响,且在控制后除一个例外外均不稳定;第三,三种线性列联表统计量在代数上不可分离;能力控制后,唯一稳定的残余项是微弱的成对共同失败关联——更多共同错误对应更低增益,该方向稳健但幅度依赖配置。将严格多样性、分歧度与双误视为独立预测因子的联合回归因构造原因秩亏。
原文摘要 · Abstract (English)
Majority voting over LLMs is widely assumed to benefit from diversity, and diversity measures are used to choose which models to combine. We ask whether five such measures track diversity or mainly re-express capability, auditing them as predictors of majority-vote gain over the best member across 31,900 subsets of 30 LLMs on MMLU-Pro (29 on TruthfulQA) under explicit capability controls. Three findings emerge. First, latent complementarity is ubiquitous: oracle gain is positive in 100% of subsets, yet simple voting beats the strongest member in only 9.98% of all canonical size-3 subsets (18.71% with held-out best selection); the pooled size-2-4 rate is 1.27%, partly reflecting deterministic even-size voting behavior. Second, a joint-correctness proxy (strict diversity) is nearly collinear with one minus mean accuracy (size-3 Spearman rho = +0.991 / +0.988); raw diversity-gain associations are strongly capability-entangled and, with one exception, unstable under control. Third, three linear contingency-table statistics are algebraically non-separable; after capability control, the empirically stable remainder is a modest residual pairwise co-failure association in which more shared error corresponds to lower gain. This direction is robust, but its magnitude is configuration-dependent. Joint rawspace linear regressions treating strict diversity, disagreement, and double-fault as independent predictors are rank-deficient by construction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。