提出量化大模型思想多样性带来的性能提升规律,可精准预测集成效果。
Quantifying Diversity of Thought: A Predictive Law of Weighted LLM Ensemble Lift
- 从理论推导出集成提升的分解公式,提炼出可预测性能的关键指标。
- 在超76万次推理中验证,预测准确率相关性达0.84以上,实测提升拟合度超0.96。
- 适用于多领域评测,尤其适合需要多模型协作的复杂任务如网络安全分析。
本文通过实验验证了一条关于大语言模型(LLM)集成中思想多样性带来性能提升的数学规律。基于第一性原理,我们精确分解了集成提升为‘救援’与‘损害’两部分,推导出简洁的预测启发式方法。从中提取出能预判集成表现的核心指标:经准确性调整的正确性相关系数ϕ_adj,以及配对模型的准确率差距和集体准确率。该规律在10个开源模型、两个研究生级科学基准测试及一个新型代理式网络安全基准上进行了验证,后者包含23,520次带放弃选项的多轮工具使用测评。在SuperGPQA上以40:60投票比例校准后,该启发式方法在训练集上的斯皮尔曼相关系数ρ=0.84;系数固定后,迁移至未参与校准的两个数据集(GPQA Diamond:ρ=0.51,取证任务:ρ=0.84),表现稳定。实测交换质量(swap mass)与实际提升的拟合度始终≥R²=0.96。原始ϕ值几乎无预测力(R²≤0.09),而ϕ_adj显著更优(SuperGPQA上R²=0.67),结合三者的启发式方法是跨数据集最稳定的预池化预测器。
原文摘要 · Abstract (English)
This paper provides an experimentally verified formal law for calculating the uplift that diversity of thought provides in Large Language Model (LLM) ensembles. From first principles, we derive an exact decomposition of LLM ensemble lift into rescue and damage masses, which yields a compact heuristic for calculating uplift. From this we extract the metrics which predict ensemble performance: an accuracy-adjusted correctness correlation, $ϕ_{\mathrm{adj}}$, together with the accuracy gap and collective accuracy of the pair. We test the law on 767,520 inferences from ten open-weight models over two graduate-level science benchmarks, together with a novel agentic cybersecurity benchmark in which each model conducts digital-forensics investigations by multi-turn tool use in a network-isolated sandbox (23,520 graded trials including abstentions); all votes are released openly. Calibrated once on SuperGPQA at a 40:60 vote split, the heuristic predicts lift on the calibration set with Spearman's $ρ=0.84$ and, with its coefficients frozen, transfers to two datasets never used in calibration ($ρ=0.51$ on GPQA Diamond and $0.84$ on the forensic tasks), whilst the measured swap mass tracks realised lift with $R^2\ge 0.96$ throughout. Raw $ϕ$ has almost no predictive power ($R^2\le 0.09$ throughout); the accuracy-adjusted $ϕ_{\mathrm{adj}}$ is markedly superior ($R^2=0.67$ on SuperGPQA), and the heuristic combining these metrics is the most stable pre-pooling predictor across the three datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。