大模型在前瞻性预测上超越人类,准确率高达0.74
The Strategic Foresight of LLMs: Evidence from a Fully Prospective Venture Tournament
- 用未见过的创业项目实时预测募资成功率
- 最佳模型预测相关性达0.74,远超人类0.45上限
- 适合关注AI决策能力的研究者与投资人
人工智能能否在战略前瞻能力上超越人类——即在不确定、高风险结果发生前做出准确判断?我们通过一场完全前瞻的预测竞赛来回答此问题,使用尚未训练过的30个美国科技初创项目(均在模型训练截止后启动)进行实时评估。多个前沿及开源大语言模型(LLMs)完成了870次配对比较,生成了完整的募资成功预测排名。我们将这些预测与通过Prolific招募的346名资深管理者及三名受控条件下训练的MBA投资者对比。结果显示:人类评估者的实际结果相关性仅为0.04至0.45,而多个前沿模型超过0.60,其中表现最优的Gemini 2.5 Pro达到0.74,几乎能正确排序五组中四组的项目。该差异在多种指标和鲁棒性检验中持续存在。无论是群体智慧集成还是人机混合团队,均未超越表现最佳的独立模型。
原文摘要 · Abstract (English)
Can artificial intelligence outperform humans at strategic foresight -- the capacity to form accurate judgments about uncertain, high-stakes outcomes before they unfold? We address this question through a fully prospective prediction tournament using live Kickstarter crowdfunding projects. Thirty U.S.-based technology ventures, launched after the training cutoffs of all models studied, were evaluated while fundraising remained in progress and outcomes were unknown. A diverse suite of frontier and open-weight large language models (LLMs) completed 870 pairwise comparisons, producing complete rankings of predicted fundraising success. We benchmarked these forecasts against 346 experienced managers recruited via Prolific and three MBA-trained investors working under monitored conditions. The results are striking: human evaluators achieved rank correlations with actual outcomes between 0.04 and 0.45, while several frontier LLMs exceeded 0.60, with the best (Gemini 2.5 Pro) reaching 0.74 -- correctly ordering nearly four of every five venture pairs. These differences persist across multiple performance metrics and robustness checks. Neither wisdom-of-the-crowd ensembles nor human-AI hybrid teams outperformed the best standalone model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。