arXiv:2411.00856cs.LGcs.AI2024-11被引 18

用大模型分析财报、行情和新闻,自动给出股票评级。

AI in Investment Analysis: LLMs for Equity Stock Ratings

  • 用大模型融合财务、市场和新闻数据生成多时长股票评级
  • 结合基本面数据时,预测准确率显著优于传统方法
  • 用情感分数替代详细新闻可省资源,不丢性能

投资分析是金融服务行业核心。随着大型语言模型(LLMs)的快速发展,其在提升股票评级流程中的潜力日益显现。传统评级依赖分析师经验,面临数据过载、文件不一致和市场反应滞后等问题。本文利用从2022年1月到2024年6月的财务、市场及新闻数据,基于GPT-4-32k(v0613,训练截止于2021年9月)构建多时域股票评级模型。结果表明,在前瞻性收益评估中,该基准方法优于传统方法,尤其在引入基本面数据时表现更佳;新闻数据有助于短期表现,但以情感得分替代详细摘要可降低令牌消耗而性能不变;在多数情况下,完全剔除新闻数据反而能提升性能,减少偏差。研究证明,大模型可有效处理大规模多模态金融数据,在股票评级任务中展现出强大能力。本工作提供了一个可复现、高效的评级框架,为传统方法提供低成本替代方案。未来将拓展至更长周期,整合更多样数据并使用更新模型以获取更深入洞察。

原文摘要 · Abstract (English)

Investment Analysis is a cornerstone of the Financial Services industry. The rapid integration of advanced machine learning techniques, particularly Large Language Models (LLMs), offers opportunities to enhance the equity rating process. This paper explores the application of LLMs to generate multi-horizon stock ratings by ingesting diverse datasets. Traditional stock rating methods rely heavily on the expertise of financial analysts, and face several challenges such as data overload, inconsistencies in filings, and delayed reactions to market events. Our study addresses these issues by leveraging LLMs to improve the accuracy and consistency of stock ratings. Additionally, we assess the efficacy of using different data modalities with LLMs for the financial domain. We utilize varied datasets comprising fundamental financial, market, and news data from January 2022 to June 2024, along with GPT-4-32k (v0613) (with a training cutoff in Sep. 2021 to prevent information leakage). Our results show that our benchmark method outperforms traditional stock rating methods when assessed by forward returns, specially when incorporating financial fundamentals. While integrating news data improves short-term performance, substituting detailed news summaries with sentiment scores reduces token use without loss of performance. In many cases, omitting news data entirely enhances performance by reducing bias. Our research shows that LLMs can be leveraged to effectively utilize large amounts of multimodal financial data, as showcased by their effectiveness at the stock rating prediction task. Our work provides a reproducible and efficient framework for generating accurate stock ratings, serving as a cost-effective alternative to traditional methods. Future work will extend to longer timeframes, incorporate diverse data, and utilize newer models for enhanced insights.

大模型股票评级金融AI多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。