用多样模型组合提升AI预测准确率,而非堆叠同类模型。
Diversity is the Strength of the AI Crowd

- 挑选预测误差互补的模型组合,而非仅选高准确率模型。
- 在Metaculus基准上,多样化模型使预测准确率显著提升。
- 像Grok 4这样的模型因低相关性贡献更大,适合做核心成员。
顶尖AI预测系统在预测未来事件上已接近超预测者水平,但仍主要依赖通用大模型(LLMs)结合特定上下文获取与框架构建。我们研究如何通过集成优化这一方法:给定固定数量的样本,应如何组合现成模型的预测以实现最高精度?在Metaculus AI基准的二元问题上,我们发现个体准确率并非关键:许多前沿大模型的预测高度相关,导致相同或相似模型的额外输入价值有限。相反,最强的集成方案是结合高准确率且预测差异大的模型,例如 exttt{Grok 4} 因其预测与其他前沿模型相关性较低,表现出显著优势。结果表明,AI群体的力量不在于盲目增加预测数量,而在于融合具有互补误差的模型,推动预测系统同时优化模型质量与多样性。
原文摘要 · Abstract (English)
Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combined with forecasting-specific context gathering and scaffolding. We study how to improve this recipe through ensembling: given a fixed number of samples, which off-the-shelf model forecasts should be combined to maximize accuracy? On binary questions from the Metaculus AI Benchmark, we find that individual accuracy is not enough: many frontier LLMs make highly correlated predictions, limiting the value of additional forecasts from the same or similar models. Instead, the strongest ensembles combine accurate but diverse forecasters, with models such as \model{Grok 4} contributing disproportionately because their predictions are less correlated with other frontier LLMs. These results suggest that the strength of the AI crowd comes not from sampling more forecasts indiscriminately, but from combining forecasts across models with complementary errors, motivating forecasting systems that explicitly optimize for both model quality and diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。