arXiv:2512.22625cs.AIcs.MA2025-12被引 1

让大模型互相讨论,能提升预测准确率。

The Wisdom of Deliberating AI Crowds: Does Deliberation Improve LLM-Based Forecasting?

  • 让不同大模型互相审阅预测结果再更新
  • 共享信息时模型组准确率提升4%
  • 适合想提高预测质量的研究者

结构化讨论已被证明能提升人类预测者的性能。本研究探究类似干预——即允许大语言模型在更新前相互审阅预测结果——是否能提升GPT-5、Claude Sonnet 4.5和Gemini Pro 2.5的预测准确性。基于Metaculus Q2 2025 AI预测竞赛中的202个已解决的二元问题,评估了四种情景:(1) 多样模型+分散信息,(2) 多样模型+共享信息,(3) 同质模型+分散信息,(4) 同质模型+共享信息。结果显示,在情景(2)中,该干预显著提升准确率,对数损失降低0.020(相对减少约4%),统计显著(p=0.017)。然而,当同质模型组(三个相同模型实例)进行同样过程时未见收益。意外的是,提供额外上下文信息并未提升预测准确率,限制了对信息整合机制的分析。研究结果表明,讨论可能是一种可行的大模型预测优化策略。

原文摘要 · Abstract (English)

Structured deliberation has been found to improve the performance of human forecasters. This study investigates whether a similar intervention, i.e. allowing LLMs to review each other's forecasts before updating, can improve accuracy in large language models (GPT-5, Claude Sonnet 4.5, Gemini Pro 2.5). Using 202 resolved binary questions from the Metaculus Q2 2025 AI Forecasting Tournament, accuracy was assessed across four scenarios: (1) diverse models with distributed information, (2) diverse models with shared information, (3) homogeneous models with distributed information, and (4) homogeneous models with shared information. Results show that the intervention significantly improves accuracy in scenario (2), reducing Log Loss by 0.020 or about 4 percent in relative terms (p = 0.017). However, when homogeneous groups (three instances of the same model) engaged in the same process, no benefit was observed. Unexpectedly, providing LLMs with additional contextual information did not improve forecast accuracy, limiting our ability to study information pooling as a mechanism. Our findings suggest that deliberation may be a viable strategy for improving LLM forecasting.

大模型预测集体智慧模型协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。