检测大模型经济预测中的提前泄露信息问题。
Detecting Lookahead Bias in LLM Forecasts
- 用日期查询法估算模型提前知晓结果的概率。
- 训练数据截止后,该概率几乎归零,但预测力仍强。
- 适合评估金融预测类LLM结果的可靠性。
我们提出一种统计方法,用于检测大型语言模型(LLMs)在经济预测中是否存在展望偏差。通过仅以日期作为查询条件,对特定公司-日期组合估计模型内化实际结果的可能性,这一指标称为展望倾向(LAP)。LAP在样本期内显著为正,而在训练数据截止后几乎降为零。我们证明,若LAP与模型预测值在准确性回归中存在正向交互作用,则表明预测结果受到展望偏差污染。该方法应用于两个任务:新闻标题预测股票收益,以及财报电话会议记录预测资本支出。在高LAP的公司-日期组合上,模型预测能力明显增强,且该交互效应在训练数据截止后的样本中不再显著。该测试提供了一种低成本、高效的诊断工具,用于评估LLM生成预测的有效性与可靠性。
原文摘要 · Abstract (English)
We develop a statistical procedure to detect lookahead bias in economic forecasts generated by large language models (LLMs). Using a date-only recall query for a firm-date pair, we estimate the probability that the LLM has internalized information about the realized outcome, a statistic we term Lookahead Propensity (LAP). LAP is materially positive throughout the in-sample period and collapses essentially to zero right after the training-data cutoff. We show that a positive interaction between LAP and the LLM forecast in an accuracy regression indicates lookahead-bias contamination, and apply the test to two forecasting tasks: news headlines predicting stock returns and earnings call transcripts predicting capital expenditures. In both applications, the LLM forecast's predictive power is amplified on high-LAP firm-date pairs, and the interaction loses significance on post-training-cutoff samples. Our test provides a cost-efficient, diagnostic tool for assessing the validity and reliability of LLM-generated forecasts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。