arXiv:2507.01970q-fin.STcs.LG2025-07被引 2

用新闻情绪向量提升股市预测,效果至少提高40%

News Sentiment Embeddings for Stock Price Forecasting

  • 用OpenAI嵌入模型将华尔街日报标题转为向量,结合PCA提取关键特征
  • 加入美元指数等经济数据后,模型在日度股价预测上性能提升超40%
  • 适合对量化金融、NLP与金融市场交叉研究感兴趣的读者

本文探讨如何利用新闻标题数据预测股票价格,目标股票为追踪美国前500家上市企业表现的SPDR S&P 500 ETF Trust(SPY)。核心方法是使用华尔街日报(WSJ)的新闻标题,通过基于OpenAI的文本嵌入模型生成每条标题的向量编码,并采用主成分分析(PCA)提取关键特征。研究重点在于捕捉新闻对股价的时变与非时变、微妙影响,同时处理潜在滞后效应和市场噪声。为提升模型表现,引入了美元指数(DXY)和国债利率等金融经济数据。共训练超过390个机器学习推理模型,初步结果显示,加入新闻标题嵌入后,股价预测性能相比未使用该数据的模型提升至少40%。

原文摘要 · Abstract (English)

This paper will discuss how headline data can be used to predict stock prices. The stock price in question is the SPDR S&P 500 ETF Trust, also known as SPY that tracks the performance of the largest 500 publicly traded corporations in the United States. A key focus is to use news headlines from the Wall Street Journal (WSJ) to predict the movement of stock prices on a daily timescale with OpenAI-based text embedding models used to create vector encodings of each headline with principal component analysis (PCA) to exact the key features. The challenge of this work is to capture the time-dependent and time-independent, nuanced impacts of news on stock prices while handling potential lag effects and market noise. Financial and economic data were collected to improve model performance; such sources include the U.S. Dollar Index (DXY) and Treasury Interest Yields. Over 390 machine-learning inference models were trained. The preliminary results show that headline data embeddings greatly benefit stock price prediction by at least 40% compared to training and optimizing a machine learning system without headline data embeddings.

股市预测新闻情绪文本嵌入量化金融

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。