arXiv:2506.00723cs.LGcs.AI2025-06被引 19

警惕大模型预测能力评估中的陷阱,避免误判其真实水平。

Pitfalls in Evaluating Language Model Forecasters

  • 发现评估中存在时间泄漏等严重问题,影响结果可信度
  • 实证显示现有评估无法可靠反映真实世界预测表现
  • 提醒研究者需建立更严谨的评测方法

大语言模型(LLMs)最近被用于预测任务,一些研究声称其表现可媲美甚至超越人类。本文指出,作为学术共同体,我们应对此类结论保持谨慎,因为评估LLM预测器面临独特挑战。我们识别出两大类问题:(1)由于多种时间泄漏形式,难以信任评估结果;(2)评估表现难以外推至真实世界预测场景。通过系统分析和对先前工作的具体案例考察,我们展示了评估缺陷如何引发对当前及未来性能宣称的担忧。我们认为,必须采用更严格的评估方法,才能自信地评估LLM的预测能力。

原文摘要 · Abstract (English)

Large language models (LLMs) have recently been applied to forecasting tasks, with some works claiming these systems match or exceed human performance. In this paper, we argue that, as a community, we should be careful about such conclusions as evaluating LLM forecasters presents unique challenges. We identify two broad categories of issues: (1) difficulty in trusting evaluation results due to many forms of temporal leakage, and (2) difficulty in extrapolating from evaluation performance to real-world forecasting. Through systematic analysis and concrete examples from prior work, we demonstrate how evaluation flaws can raise concerns about current and future performance claims. We argue that more rigorous evaluation methodologies are needed to confidently assess the forecasting abilities of LLMs.

大模型评估预测能力时间泄漏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。