测试大模型在预测任务中使用不同上下文的效果。
Analyzing the Role of Context in Forecasting with Large Language Models
- 构建600多个带新闻摘要的二元预测题数据集
- 加入新闻能显著提升预测准确率,少样本示例反而降低效果
- 大模型比小模型更优,适合自动化预测场景
本研究评估了近期语言模型(LLMs)在二元预测问题上的表现。我们首先构建了一个包含600多个二元预测问题的数据集,每个问题均配有相关新闻文章及其简洁的问题关联摘要。随后,我们探讨了不同层次上下文输入对预测性能的影响。结果表明,引入新闻文章可显著提升模型表现,而使用少量示例(few-shot)反而导致准确率下降。我们发现,大模型始终优于小模型,凸显了大型语言模型在提升自动化预测能力方面的潜力。
原文摘要 · Abstract (English)
This study evaluates the forecasting performance of recent language models (LLMs) on binary forecasting questions. We first introduce a novel dataset of over 600 binary forecasting questions, augmented with related news articles and their concise question-related summaries. We then explore the impact of input prompts with varying level of context on forecasting performance. The results indicate that incorporating news articles significantly improves performance, while using few-shot examples leads to a decline in accuracy. We find that larger models consistently outperform smaller models, highlighting the potential of LLMs in enhancing automated forecasting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。