arXiv:2508.11987cs.AIcs.LG2025-08被引 43

构建首个实时更新的未来预测评估基准,测试大模型在动态环境中的推理能力。

FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction

  • 设计实时动态基准,自动收集问题与答案,防止数据污染。
  • 评测25个模型,发现其在虚假网页和时间有效性上存在明显漏洞。
  • 适合研究智能体预测、复杂推理及动态决策的开发者与研究人员。

未来预测是大语言模型智能体的一项复杂任务,需具备高水平的分析思维、信息获取、上下文理解及不确定性下的决策能力。智能体不仅要处理海量动态信息,还需整合多源数据、权衡不确定性,并根据新趋势调整预测,如同政治、经济、金融等领域的专业人员。然而,由于实时更新与准确回答获取困难,当前尚无大规模评估基准。为此,我们提出$ extbf{FutureX}$,一个专为未来预测设计的动态、实时评估基准。FutureX 是目前规模最大、多样性最高的实时预测评估平台,支持每日自动更新,通过自动化流程实现问题与答案的采集,杜绝数据污染。我们评估了25个大模型/智能体,包括具备推理、搜索能力及外部工具集成的开源(如Deep Research Agent)与闭源模型。全面分析了智能体在动态环境中的适应性推理表现及其失败模式,揭示其对虚假网页和时间有效性的脆弱性。目标是建立一个动态、无污染的评估标准,推动大模型智能体达到专业人类分析师的复杂推理与预测水平。

原文摘要 · Abstract (English)

Future prediction is a complex task for LLM agents, requiring a high level of analytical thinking, information gathering, contextual understanding, and decision-making under uncertainty. Agents must not only gather and interpret vast amounts of dynamic information but also integrate diverse data sources, weigh uncertainties, and adapt predictions based on emerging trends, just as human experts do in fields like politics, economics, and finance. Despite its importance, no large-scale benchmark exists for evaluating agents on future prediction, largely due to challenges in handling real-time updates and retrieving timely, accurate answers. To address this, we introduce $\textbf{FutureX}$, a dynamic and live evaluation benchmark specifically designed for LLM agents performing future prediction tasks. FutureX is the largest and most diverse live benchmark for future prediction, supporting real-time daily updates and eliminating data contamination through an automated pipeline for question gathering and answer collection. We evaluate 25 LLM/agent models, including those with reasoning, search capabilities, and integration of external tools such as the open-source Deep Research Agent and closed-source Deep Research models. This comprehensive evaluation assesses agents' adaptive reasoning and performance in dynamic environments. Additionally, we provide in-depth analyses of agents' failure modes and performance pitfalls in future-oriented tasks, including the vulnerability to fake web pages and the temporal validity. Our goal is to establish a dynamic, contamination-free evaluation standard that drives the development of LLM agents capable of performing at the level of professional human analysts in complex reasoning and predictive thinking.

大模型评估未来预测动态基准智能体推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。