arXiv:2604.07355cs.LGcs.AI2026-04被引 8

让大模型真金白银炒股,测试它们在真实市场的预测能力

Prediction Arena: Benchmarking AI Models on Real-World Prediction Markets

论文配图:Prediction Arena: Benchmarking AI Models on Real-World Prediction Markets
图 1 · 摘自论文原文
  • 模型自主交易,每15-45分钟决策一次,用真实资金在Kalshi和Polymarket平台实战
  • 6个前沿模型57天实盘后收益介于-16.0%到-30.8%,但同一模型在不同平台表现差异巨大
  • 平台设计影响显著:同一模型在Polymarket平均赚6.02%,而同平台另一模型胜率超71%

我们提出Prediction Arena,一个通过让AI模型在真实预测市场中自主交易并使用真实资本来评估其预测准确性和决策能力的基准。不同于合成基准,Prediction Arena在实际交易所(Kalshi和Polymarket)上测试模型,提供无法被操纵或过拟合的真实反馈。每个模型作为独立代理,初始资金1万美元,每15-45分钟自主决策。在为期57天的纵向评估(2026年1月12日至3月9日)中,跟踪两组模型:六种前沿模型进行实盘交易(第一组,全程),四种下一代模型进行纸面交易(第二组,三天预演)。第一组在Kalshi的最终回报率为-16.0%至-30.8%。分析显示明确的性能层级:初始预测准确性和兑现正确预测的能力是主要驱动因素,而研究量与结果无相关性。在平行的Polymarket实盘交易中,第一组模型平均回报为-1.1%,远高于在Kalshi的-22.6%,其中grok-4-20-checkpoint实现71.4%的结算胜率,为所有平台和组别中的最高。gemini-3.1-pro-preview(第二组)在Kalshi未执行任何交易,但在3天内于Polymarket获得+6.02%回报,是两组中最佳表现,表明平台设计对模型成功有深远影响。除性能外,我们还分析了计算效率(令牌使用、周期时间)、结算准确率、退出模式和市场偏好,全面揭示前沿模型在真实金融压力下的行为。

原文摘要 · Abstract (English)

We introduce Prediction Arena, a benchmark for evaluating AI models' predictive accuracy and decision-making by enabling them to trade autonomously on live prediction markets with real capital. Unlike synthetic benchmarks, Prediction Arena tests models in environments where trades execute on actual exchanges (Kalshi and Polymarket), providing objective ground truth that cannot be gamed or overfitted. Each model operates as an independent agent starting with $10,000, making autonomous decisions every 15-45 minutes. Over a 57-day longitudinal evaluation (January 12 to March 9, 2026), we track two cohorts: six frontier models in live trading (Cohort 1, full period) and four next-generation models in paper trading (Cohort 2, 3-day preliminary). For Cohort 1, final Kalshi returns range from -16.0% to -30.8%. Our analysis identifies a clear performance hierarchy: initial prediction accuracy and the ability to capitalize on correct predictions are the main drivers, while research volume shows no correlation with outcomes. A striking cross-platform contrast emerges from parallel Polymarket live trading: Cohort 1 models averaged only -1.1% on Polymarket vs. -22.6% on Kalshi, with grok-4-20-checkpoint achieving a 71.4% settlement win rate - the highest across any platform or cohort. gemini-3.1-pro-preview (Cohort 2), which executed zero trades on Kalshi, achieved +6.02% on Polymarket in 3 days - the best return of any model across either cohort - demonstrating that platform design has a profound effect on which models succeed. Beyond performance, we analyze computational efficiency (token usage, cycle time), settlement accuracy, exit patterns, and market preferences, providing a comprehensive view of how frontier models behave under real financial pressure.

预测市场大模型评测实盘测试AI金融

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。