用15个大模型做集体预测,发现组合效果优于单个模型。
Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles
- 用神经网络和逻辑回归融合多个大模型输出,提升预测准确率。
- 模型组合后平均准确率比最强单个模型高35.8%。
- 适合关注模型集成与评估公平性的研究人员参考。
我们从15个大语言模型获取254个二元预测市场的概率估计,并评估经典与学习型聚合方法。学习型聚合器(多层感知机与逻辑回归)均优于所有单个模型及经典方法。逻辑回归表现与神经网络相当,表明其优势来自对多样模型输出的线性组合,而非非线性交互。对神经网络学习映射进行符号回归,发现纯模型分歧信号是帕累托前沿上最简有效公式,进一步支持该结论。训练截止时间污染是普遍干扰:在所有模型训练截止后才解决的问题子集上,前沿模型与小型本地模型的性能差距由35.8%降至8.9%,且单个模型排名仅中等稳定。即使在各模型训练截止时评估,大模型群体仍显著低于人类预测市场,表明其集体信息整合能力存在真实差距。结果表明,大模型群体可呈现智慧之群效应,但无污染评估是可靠判断的前提。
原文摘要 · Abstract (English)
The wisdom of crowds -- the finding that aggregating judgments across individuals often outperforms the best individual -- has been extensively studied with human forecasters. Whether the same phenomenon emerges when the ``crowd'' consists of large language models (LLMs) is an open question with both theoretical and practical implications. We elicited probability estimates from 15 LLMs on 254 binary prediction market questions and evaluated classical and learned aggregation methods. Learned aggregators -- a multilayer perceptron and a logistic regression -- outperformed all individual models and classical methods. The logistic regression was found to match the neural network, suggesting that the benefit of learned aggregation derives from learning a linear combination of diverse model outputs rather than from nonlinear interactions. Symbolic regression applied to the neural network's learned mapping recovered a pure model-disagreement signal as the lowest-complexity useful formula on the Pareto frontier, further supporting this interpretation. Training cutoff contamination proved a pervasive confound: the apparent capability gap between frontier cloud models and smaller local models collapsed from 35.8% to 8.9% on a clean subset of questions resolving after all models' training cutoffs, and individual model rankings showed only moderate stability. Even when the prediction market is evaluated at each model's training cutoff, LLMs remained substantially less accurate, indicating a genuine gap in collective information aggregation. These findings suggest that LLM crowds can exhibit wisdom-of-crowds effects, but that contamination-free evaluation is essential for reliable assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。