研究大模型在预测市场生命周期中的可靠性变化,发现早期和高不确定性时表现更好。
TimeSeek: Temporal Reliability of Agentic Forecasters
- 在150个受监管的二元市场中,分五个时间点评估10个前沿模型的预测表现。
- 模型在市场初期和高不确定性市场中表现最佳,临近结束时准确率显著下降。
- 网络搜索整体提升预测质量,但部分场景下反而有害,适合动态调整使用策略。
我们提出TimeSeek,一个用于研究代理型大模型预测者在预测市场生命周期中可靠性变化的基准。在150个受CFTC监管的Kalshi二元市场中,对10个前沿模型在五个时间点进行评估,包含有无网络搜索的条件,共生成15,000次预测。结果显示,模型在市场早期和高不确定性市场中最具竞争力,但在接近结算时以及强共识市场中表现大幅下降。网络搜索整体提升了所有模型的合并贝叶斯评分(BSS),但在12%的模型-时间点组合中反而造成性能下降,表明检索工具虽平均有益,但并非普遍适用。简单的双模型集成可降低误差,但未能超越市场整体表现。这些结果支持采用时间感知的评估方法与选择性依赖策略,而非单一时间快照或统一工具使用设置。
原文摘要 · Abstract (English)
We introduce TimeSeek, a benchmark for studying how the reliability of agentic LLM forecasters changes over a prediction market's lifecycle. We evaluate 10 frontier models on 150 CFTC-regulated Kalshi binary markets at five temporal checkpoints, with and without web search, for 15,000 forecasts total. Models are most competitive early in a market's life and on high-uncertainty markets, but much less competitive near resolution and on strong-consensus markets. Web search improves pooled Brier Skill Score (BSS) for every model overall, yet hurts in 12% of model-checkpoint pairs, indicating that retrieval is helpful on average but not uniformly so. Simple two-model ensembles reduce error without surpassing the market overall. These descriptive results motivate time-aware evaluation and selective-deference policies rather than a single market snapshot or a uniform tool-use setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。