arXiv:2604.14199q-fin.CPcs.AI2026-04被引 11

用真实市场数据测试大模型预测能力,发现多数模型看似自信却亏钱。

PolyBench: Benchmarking LLM Forecasting and Trading Capabilities on Live Prediction Market Data

  • 基于实时订单簿和新闻流构建多模态预测基准
  • 仅两款模型实现正收益,最高17.6%信心加权回报
  • 适合关注金融推理与真实场景评估的研究者

从实时市场信号中预测真实事件,需要融合定性新闻与定量订单簿动态的系统,并在严格时间约束下运行——现有基准无法捕捉这一挑战。我们提出 extbf{PolyBench},一个源自 Polymarket 的多模态基准,记录了 38,666 个二元预测市场在 4,997 个事件上的时点快照,同步关联每个快照的中央限价订单簿(CLOB)状态与实时新闻流。利用 PolyBench,我们评估了七款先进大语言模型(涵盖开源与闭源系列),在 2026 年 2 月 6 日至 12 日期间,于相同时间锁定的市场状态下生成 36,165 条预测。我们的多维评估框架包含方向准确率、新提出的置信度加权回报(CWR)、年化收益率(APY)及夏普比率,通过真实的订单簿执行模拟进行测算。结果揭示显著性能差异:仅有两款模型实现正财务回报——MiMo-V2-Flash 达到 17.6% CWR,Gemini-3-Flash 为 6.2% CWR,其余五款尽管普遍高置信度,却均产生亏损。这凸显了表面语言流畅性与真实市场不确定性下概率推理能力之间的鸿沟,并确立 PolyBench 作为抗污染、金融导向的未来大模型研究评估标准。数据与代码见:https://github.com/PolyBench/PolyBench。

原文摘要 · Abstract (English)

Predicting real-world events from live market signals demands systems that fuse qualitative news with quantitative order-book dynamics under strict temporal discipline -- a challenge existing benchmarks fail to capture. We present \textbf{PolyBench}, a multimodal benchmark derived from Polymarket that records point-in-time cross-sections of 38,666 binary prediction markets spanning 4,997 events, synchronously coupling each snapshot with a Central Limit Order Book (CLOB) state and a real-time news stream. Using PolyBench, we evaluate seven state-of-the-art Large Language Models -- spanning open- and closed-source families -- generating 36,165 predictions under identical, timestamp-locked market states collected between February 6 and 12, 2026. Our multidimensional framework assesses directional accuracy, our proposed Confidence-Weighted Return (CWR), Annualized Percentage Yield (APY), and Sharpe ratio via realistic order-book execution simulation. The results reveal a pronounced performance divergence: only two of seven models achieve positive financial returns -- MiMo-V2-Flash at \textbf{17.6%} CWR and Gemini-3-Flash at 6.2% CWR -- while the remaining five incur losses despite uniformly high stated confidence. These findings highlight the gap between surface-level language fluency and genuine probabilistic reasoning under live market uncertainty, and establish PolyBench as a contamination-proof, financially-grounded evaluation standard for future LLM research. Our dataset and code available at \underline{\href{https://github.com/PolyBench/PolyBench}{https://github.com/PolyBench/PolyBench}}.

大模型评估金融预测多模态基准市场模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。