arXiv:2501.05925cs.CLcs.IR2025-01被引 5

评测大模型对未来事件的预测能力,发现其潜力与局限。

Navigating Tomorrow: Reliably Assessing Large Language Models Performance on Future Event Prediction

  • 构建新闻数据集,测试大模型在三种预测场景下的表现。
  • 模型在可能性判断和反事实分析中表现较好,但受限于训练数据截止时间。
  • 适合关注大模型推理能力与未来预测应用的研究者参考。

预测未来事件在多个领域具有重要意义,如预判股市趋势、自然灾害、商业动态或政治变化,可促进早期防范并发掘新机遇。已有多种计算方法用于未来预测,包括预测分析、时间序列建模和仿真等。本研究评估了多个大语言模型(LLMs)在支持未来预测任务中的表现,这是尚未充分探索的领域。我们设计了三个评估场景:肯定性与可能性提问、推理分析、反事实推演。为此,我们基于实体类型和热度对新闻文章进行搜集与分类,覆盖模型训练截止日期前后的文章,以全面测试和比较模型性能。研究揭示了大模型在预测建模中的潜力与局限,为后续改进提供了基础。

原文摘要 · Abstract (English)

Predicting future events is an important activity with applications across multiple fields and domains. For example, the capacity to foresee stock market trends, natural disasters, business developments, or political events can facilitate early preventive measures and uncover new opportunities. Multiple diverse computational methods for attempting future predictions, including predictive analysis, time series forecasting, and simulations have been proposed. This study evaluates the performance of several large language models (LLMs) in supporting future prediction tasks, an under-explored domain. We assess the models across three scenarios: Affirmative vs. Likelihood questioning, Reasoning, and Counterfactual analysis. For this, we create a dataset1 by finding and categorizing news articles based on entity type and its popularity. We gather news articles before and after the LLMs training cutoff date in order to thoroughly test and compare model performance. Our research highlights LLMs potential and limitations in predictive modeling, providing a foundation for future improvements.

未来预测大模型评估推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。