arXiv:2605.16358cs.LGcs.AI2026-05

首个动态更新的事件增强预测基准,评估大模型在真实世界中的预测能力。

LEAF: A Living Benchmark for Event-Augmented Forecasting

论文配图:LEAF: A Living Benchmark for Event-Augmented Forecasting
图 1 · 摘自论文原文
  • 用递归检索代理+双代理交叉验证生成辅助信息
  • 大模型能利用复杂事件信号提升股票预测性能
  • 适合研究事件驱动预测与大模型评估的学者

大语言模型(LLMs)在预测任务中应用日益广泛。为评估其预测能力并避免预训练数据污染,已有若干动态基准被提出。然而,现有基准或因数据稀缺缺乏多维事件,或聚焦于相对封闭环境。为评估大模型在复杂真实场景下的预测能力,我们提出LEAF——首个面向事件增强预测任务的动态基准,涵盖未来事件概率、趋势及时间序列预测。LEAF采用递归检索代理系统与双代理交叉验证机制,提供全面且相关的辅助文本。评估主流专有与开源大模型发现,这些模型能有效利用复杂事件中的信号提升预测表现。在股票领域,模型对自身判断更可预测的股票表现出更高性能,且事件与目标股票强相关。LEAF为事件驱动预测任务提供了必要且持续更新的测试平台,可持续追踪与推动该领域进展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly applied to forecasting. To evaluate this capability while mitigating pre-training data contamination, several living benchmarks have been proposed. However, existing benchmarks either lack the multidimensional events essential for accurate forecasting due to data scarcity, or focus on relatively closed environments. To assess the predictive capabilities of LLMs in complex, real-world scenarios, we propose LEAF, the first living benchmark for event-augmented forecasting tasks, including future event probabilities, trend and time series forecasting. LEAF utilizes a recursive retrieval agent system paired with dual-agent cross-validation to provide comprehensive and relevant auxiliary text for forecasting. Evaluating state-of-the-art proprietary and open-weight LLMs, we find that these models can leverage signals extracted from complex events to enhance predictive performance. In the stock domain, we find that LLMs achieve better performance on equities they confidently identify as more predictable. Furthermore, the events demonstrate a strong correlation with the target equities. To this end, LEAF provides a necessary, dynamically updating testbed to continuously track and drive progress in event-driven forecasting tasks.

事件预测大模型评估动态基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。