用大模型预测2030年出现通用人工智能的可能性,发现其判断与专家意见高度吻合。
AI Predicts AGI: Leveraging AGI Forecasting and Peer Review to Explore LLMs' Complex Reasoning Capabilities
- 让16个顶级大模型评估AGI在2030年前出现的概率,方法为自动化同行评审。
- 模型预测范围从3%到47.6%,中位数12.5%,与专家调查结果接近。
- 提出新基准评估模型在复杂推断任务中的表现,适合关注AI未来评估的研究者。
我们让16个最先进的大语言模型(LLMs)评估通用人工智能(AGI)在2030年前出现的可能性。为评估这些预测质量,我们设计了自动化同行评审系统(LLM-PR)。模型预测值差异显著,介于3%(Reka-Core)至47.6%(GPT-4o)之间,中位数为12.5%。该结果与近期专家调查中10%的预测概率高度一致,表明大模型在复杂、推测性场景中具有参考价值。LLM-PR展现出高可靠性,组内相关系数(ICC)达0.79,评分一致性良好。其中,Pplx-70b-online表现最优,Gemini-1.5-pro-api最差。与外部基准(如LMSYS Chatbot Arena)对比发现,模型排名在不同评估方式下保持一致,暗示现有基准可能无法涵盖AGI预测所需能力。我们进一步基于外部基准引入加权策略,优化模型预测与人类专家意见的一致性,并构建了新的“AGI基准”以突出模型在相关任务中的差异表现。研究揭示了大模型在跨学科推测性任务中的能力,强调需发展新型评估框架来应对真实世界中的不确定性挑战。
原文摘要 · Abstract (English)
We tasked 16 state-of-the-art large language models (LLMs) with estimating the likelihood of Artificial General Intelligence (AGI) emerging by 2030. To assess the quality of these forecasts, we implemented an automated peer review process (LLM-PR). The LLMs' estimates varied widely, ranging from 3% (Reka- Core) to 47.6% (GPT-4o), with a median of 12.5%. These estimates closely align with a recent expert survey that projected a 10% likelihood of AGI by 2027, underscoring the relevance of LLMs in forecasting complex, speculative scenarios. The LLM-PR process demonstrated strong reliability, evidenced by a high Intraclass Correlation Coefficient (ICC = 0.79), reflecting notable consistency in scoring across the models. Among the models, Pplx-70b-online emerged as the top performer, while Gemini-1.5-pro-api ranked the lowest. A cross-comparison with external benchmarks, such as LMSYS Chatbot Arena, revealed that LLM rankings remained consistent across different evaluation methods, suggesting that existing benchmarks may not encapsulate some of the skills relevant for AGI prediction. We further explored the use of weighting schemes based on external benchmarks, optimizing the alignment of LLMs' predictions with human expert forecasts. This analysis led to the development of a new, 'AGI benchmark' designed to highlight performance differences in AGI-related tasks. Our findings offer insights into LLMs' capabilities in speculative, interdisciplinary forecasting tasks and emphasize the growing need for innovative evaluation frameworks for assessing AI performance in complex, uncertain real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。