arXiv:2606.21013cs.AIcs.LG2026-06被引 1

用时间机器框架加速大模型未来事件预测评估,效果媲美实时测试

Agentic Time Machine as an Infrastructure for Future-Event Forecasting

论文配图:Agentic Time Machine as an Infrastructure for Future-Event Forecasting
图 1 · 摘自论文原文
  • 构建可回溯网络状态的代理时间机器,实现快速仿真评估
  • 多智能体框架并行分析,预测准确率超越多个基线模型
  • 适合需高效验证预测能力的AI研究者与量化团队使用

大语言模型代理在选举、货币政策和金融市场等未来事件预测中面临关键挑战。现有评估方法在效率与环境真实度间存在根本权衡:实时评测反馈慢,而历史回放通常依赖静态数据,缺乏实际部署的动态性。为此,我们提出代理时间机器(Agentic Time Machine, TM),通过过滤截断后内容,近似重建任意过去时间点的网络状态。基于此基础设施,我们设计了一个规划-求解-聚合的多智能体框架,将问题分解为多元分析视角,平行收集证据并融合结果生成预测。实验表明,TM下的离线评分与实时FutureX评分高度相关,验证其作为快速可靠评估沙盒的有效性。在FutureX-Past与Polymarket上,该框架在封闭书、工具增强及自洽性基线中表现最优;在官方FutureX实时排行榜中,连续四周平均排名最高,五月第一周位列第一;截至6月17日,八周总榜排名第一。

原文摘要 · Abstract (English)

Forecasting future events is a critical challenge for large language model (LLM) agents, spanning domains from elections and monetary policy to financial markets. However, evaluating progress on this task presents a fundamental trade-off between efficiency and environment fidelity. While live evaluation benchmarks suffer from an inherently slow feedback loop, existing retrospective replays typically restrict agents to static, pre-frozen databases that sacrifice the environmental realism of actual deployments. To tackle this issue, we introduce Agentic Time Machine (TM), an infrastructure that approximately reconstructs the web state at any chosen past time by filtering post-cutoff content. Leveraging this evaluation infrastructure, we further propose a planner-solver-aggregator multi-agent framework that breaks each question into diverse analytical angles, gathers evidence in parallel, and combines the results into a single forecast. Experiments show that offline scores under TM correlate strongly with live FutureX scores, validating that TM offers a fast and reliable sandbox for forecasting-agent evaluation. On FutureX-Past and Polymarket evaluated under TM, our framework achieves the highest score among strong closed-book, tool-augmented, and self-consistency baselines. On the official FutureX live leaderboard, our system achieves the best average rank over four consecutive weeks, including 1st place in May Week 1. As of June 17, it also ranks 1st on FutureX's official eight-week overall leaderboard.

未来预测多智能体评估框架时间机器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。