首个链上预测市场基准,用真实数据评估AI预测能力。
Foresight Arena: An On-Chain Benchmark for Evaluating AI Forecasting Agents
- 通过智能合约在Polygon链上实现无信任的预测提交与结算
- 采用布里尔分数和新型阿尔法得分,准确区分预测能力与市场时机
- 可检测0.02以上真实预测优势,适合研究可信AI预测的学者
评估AI预测代理的真实能力需要抗过拟合、去中心化且激励相容的环境。现有基准或依赖易污染的静态数据集,或使用交易盈亏(PnL)——该指标混淆了预测准确性与交易时机、仓位大小和风险偏好。本文提出Foresight Arena,首个无需许可、基于链上的AI预测代理评测基准,用于真实世界预测市场。代理通过提交-揭示协议,在Polygon PoS链上以Solidity智能合约执行对二元Polymarket市场的概率预测;结果通过Gnosis条件代币框架无信任结算。性能采用布里尔分数(Brier Score)和新型阿尔法得分(Alpha Score)衡量,二者均为合理评分规则,鼓励诚实的概率报告,并分离出相对于市场共识的预测优势。我们进行了形式化分析:给出了单市场阿尔法得分的闭式方差表达式,建立了与穆尔菲经典布里尔分解的联系,并进行功效分析,刻画了可靠区分不同技能水平代理所需的轮数。结果显示,在80%功效下,检测真实优势α* = 0.02需约350个已决断二元预测(50轮×7个市场),而α* = 0.01则需四倍数量。我们还通过校准至文献报道布里尔分数范围的确定性模拟研究,验证了穆尔菲分解可有效区分校准良好代理与仅追踪市场的代理。部署后的实时结果将在后续版本中公布。所有智能合约与评估基础设施均开源。
原文摘要 · Abstract (English)
Evaluating the true forecasting ability of AI agents requires environments that are resistant to environments resistant to overfitting, free from centralized trust, and grounded in incentive-compatible scoring. Existing benchmarks either rely on static datasets vulnerable to training-data contamination, or measure trading PnL -- a metric conflating predictive accuracy with timing, sizing, and risk appetite. We introduce Foresight Arena, the first permissionless, on-chain benchmark for evaluating AI forecasting agents on real-world prediction markets. Agents submit probabilistic forecasts on binary Polymarket markets via a commit-reveal protocol enforced by Solidity smart contracts on Polygon PoS; outcomes are resolved trustlessly through the Gnosis Conditional Token Framework. Performance is measured by the Brier Score and a novel Alpha Score -- proper scoring rules that incentivize honest probability reporting and isolate predictive edge over market consensus. We provide a formal analysis: closed-form variance for per-market Alpha, the connection to Murphy's classical Brier decomposition, and a power analysis characterizing the number of rounds required to reliably distinguish agents of different skill levels. We show that detecting a true edge of $α^* = 0.02$ at 80% power requires approximately 350 resolved binary predictions (50 rounds of 7 markets), while $α^* = 0.01$ requires four times more. We complement these analytical results with a deterministic, seed-controlled simulation study calibrated to literature-reported Brier-score ranges, illustrating how Murphy decomposition distinguishes well-calibrated agents from market-tracking agents that fail through reduced resolution. Live results from the deployed benchmark will be reported in a future revision. All smart contracts and evaluation infrastructure are open-source.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。