提出可推理的未来事件预测基准,用因果干预评估预测可信度。
PROPHET: An Inferable Future Forecasting Benchmark with Causal Intervened Likelihood Estimation
- 通过因果干预概率衡量问题是否可从新闻中合理推断
- 构建首个保证可推理性的未来预测数据集,筛选后保留300+有效问题
- 适合研究可解释预测、因果推理与大模型可信生成的学者
基于网络新闻预测未来事件是人工智能的终极目标之一。近年来,基于大语言模型的系统在事件预测方面展现出巨大潜力,研究社区对此高度关注。现有多个基准将事件预测任务形式化为检索增强生成与推理任务,每个问题通过从网络下载的相关新闻文章回答。然而,这些基准未考虑问题是否具备有效的支持理由,部分问题本身不可推理。为此,我们提出新基准PROPHET,包含可推理的预测问题及其相关新闻。为确保可推理性,我们提出因果干预似然(CIL)这一统计指标,通过因果推断评估可推理性。首先收集近期趋势预测问题,再用CIL筛选,构建可推理的未来预测基准。通过大量实验,我们验证了CIL的有效性,并深入分析其在预测中的应用。随后在PROPHET上评估多种代表性预测方法,结果揭示了未来方向的关键洞察。
原文摘要 · Abstract (English)
Predicting future events based on news on the Web stands as one of the ultimate aspirations of artificial intelligence. Recent advances in large language model (LLM)-based systems have shown remarkable potential in forecasting future events, thereby garnering significant interest in the research community. Currently, several benchmarks have been established to evaluate the forecasting capabilities by formalizing the event prediction as a retrieval-augmented generation (RAG)-and-reasoning task. In these benchmarks, each prediction question is answered with relevant retrieved news articles downloaded from the Web. However, because there is no consideration of whether the questions can be supported by valid or sufficient supporting rationales, some of the questions in these benchmarks may be inherently noninferable. To address this issue, we introduce a new benchmark, PROPHET, which comprises inferable forecasting questions paired with relevant news for retrieval. To ensure the inferability of the benchmark, we propose Causal Intervened Likelihood (CIL), a statistical measure that assesses inferability through causal inference. In constructing this benchmark, we first collected recent trend forecasting questions, and then filtered the data using CIL resulting in an inferable benchmark for future forecasting. Through extensive experiments, we first demonstrate the validity of CIL and in-depth investigations into future forecasting with the aid of CIL. Subsequently, we evaluate several representative prediction methods on PROPHET. The overall results draws valuable insights for task of future directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。