用实时数据评测大模型炒股能力,发现其易受假新闻影响
PriceSeer: Evaluating Large Language Models in Real-Time Stock Prediction
- 构建实时动态的股票预测基准,含110只美股249个历史点
- 大模型在短期预测中表现不错,长期预测效果差且易被假新闻干扰
- 适合研究金融AI、大模型泛化能力的学者与开发者参考
股票预测是与人们投资活动紧密相关的动态实时任务,已有广泛研究。当前大型语言模型(LLMs)在多个领域展现出卓越潜力,通过先进推理和上下文理解实现专家级表现。本文提出PriceSeer,一个专为大模型股票预测设计的实时、动态、数据无污染基准。该基准包含110只美国股票,覆盖11个行业,每只股票有249个历史数据点。系统支持内部与外部信息扩展,让模型接收财务指标、新闻及虚假新闻以进行股价预测。我们评估了六种前沿大模型在不同预测周期下的表现,验证其生成投资策略的潜力。同时分析了模型在长期预测中的不足,包括对假新闻的脆弱性及特定行业的表现偏差。代码与评估数据将开源于https://github.com/BobLiang2113/PriceSeer。
原文摘要 · Abstract (English)
Stock prediction, a subject closely related to people's investment activities in fully dynamic and live environments, has been widely studied. Current large language models (LLMs) have shown remarkable potential in various domains, exhibiting expert-level performance through advanced reasoning and contextual understanding. In this paper, we introduce PriceSeer, a live, dynamic, and data-uncontaminated benchmark specifically designed for LLMs performing stock prediction tasks. Specifically, PriceSeer includes 110 U.S. stocks from 11 industrial sectors, with each containing 249 historical data points. Our benchmark implements both internal and external information expansion, where LLMs receive extra financial indicators, news, and fake news to perform stock price prediction. We evaluate six cutting-edge LLMs under different prediction horizons, demonstrating their potential in generating investment strategies after obtaining accurate price predictions for different sectors. Additionally, we provide analyses of LLMs' suboptimal performance in long-term predictions, including the vulnerability to fake news and specific industries. The code and evaluation data will be open-sourced at https://github.com/BobLiang2113/PriceSeer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。