用真实世界事件回放测试AI代理的长期适应能力
FutureSim: Replaying World Events to Evaluate Adaptive Agents

- 构建时间线回放系统,让代理在真实新闻流中预测未来事件
- 最佳代理三月预测准确率仅25%,多数表现不如随机猜测
- 适合研究长时序适应、推理与不确定性处理的前沿方向
AI代理正被部署于动态开放环境,需随新信息持续适应。为高效评估此类能力,我们提出构建基于真实事件回放的地面实况模拟系统。建立FutureSim,让代理在已知截止时间后,面对按时间顺序重现的真实新闻和问题解答进行预测。我们在原生环境中评估前沿代理,测试其对2026年1月至3月期间世界事件的预测能力。结果揭示显著能力差异:最优代理准确率为25%,许多代理的布里尔技能得分甚至低于不做预测。通过细致消融实验,证明FutureSim能真实反映长时程测试时自适应、搜索、记忆与不确定性推理等新兴研究方向。整体而言,该基准设计有望推动对真实世界长期适应能力的评估。
原文摘要 · Abstract (English)
AI agents are being increasingly deployed in dynamic, open-ended environments that require adapting to new information as it arrives. To efficiently measure this capability for realistic use-cases, we propose building grounded simulations that replay real-world events in the order they occurred. We build FutureSim, where agents forecast world events beyond their knowledge cutoff while interacting with a chronological replay of the world: real news articles arriving and questions resolving over the simulated period. We evaluate frontier agents in their native harness, testing their ability to predict world events over a three-month period from January to March 2026. FutureSim reveals a clear separation in their capabilities, with the best agent's accuracy being 25%, and many having worse Brier skill score than making no prediction at all. Through careful ablations, we show how FutureSim offers a realistic setting to study emerging research directions like long-horizon test-time adaptation, search, memory, and reasoning about uncertainty. Overall, we hope our benchmark design paves the way to measure AI progress on open-ended adaptation spanning long time-horizons in the real world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。