arXiv:2503.08696q-fin.STcs.LG2025-03被引 4

融合新闻文本与股价数据,显著提升俄股预测精度。

Multimodal Stock Price Prediction: A Case Study of the Russian Securities Market

  • 用LSTM同时处理蜡烛图数据和新闻文本,实现多模态预测。
  • 加入新闻文本后,价格预测误差(MAPE)降低55%。
  • 适合金融量化研究者及关注市场情绪的投资者参考。

传统资产价格预测主要依赖价格序列、交易量、订单簿数据和技术指标等数值信息。然而新闻流在价格形成中起关键作用,因此结合文本与数值数据的多模态方法对提升预测准确性具有重要意义。本文针对莫斯科交易所176只俄罗斯股票,构建了包含79,555篇俄语财经新闻的文章的多模态数据集,结合蜡烛图时间序列与新闻文本数据进行股价预测。文本采用RuBERT和Vikhr-Qwen2.5-0.5b-Instruct预训练模型处理,时间序列与向量化文本由LSTM网络建模。实验对比了单模态(仅时间序列)与双模态模型,以及不同文本向量聚合方法。评估指标包括方向预测准确率(Accuracy)和平均绝对百分比误差(MAPE)。结果表明,引入文本模态使MAPE降低55%。该多模态数据集可为金融领域语言模型的适配提供支持。未来工作包括优化新闻的时间窗口、情感分析与发布时间顺序等参数。

原文摘要 · Abstract (English)

Classical asset price forecasting methods primarily rely on numerical data, such as price time series, trading volumes, limit order book data, and technical analysis indicators. However, the news flow plays a significant role in price formation, making the development of multimodal approaches that combine textual and numerical data for improved prediction accuracy highly relevant. This paper addresses the problem of forecasting financial asset prices using the multimodal approach that combines candlestick time series and textual news flow data. A unique dataset was collected for the study, which includes time series for 176 Russian stocks traded on the Moscow Exchange and 79,555 financial news articles in Russian. For processing textual data, pre-trained models RuBERT and Vikhr-Qwen2.5-0.5b-Instruct (a large language model) were used, while time series and vectorized text data were processed using an LSTM recurrent neural network. The experiments compared models based on a single modality (time series only) and two modalities, as well as various methods for aggregating text vector representations. Prediction quality was estimated using two key metrics: Accuracy (direction of price movement prediction: up or down) and Mean Absolute Percentage Error (MAPE), which measures the deviation of the predicted price from the true price. The experiments showed that incorporating textual modality reduced the MAPE value by 55%. The resulting multimodal dataset holds value for the further adaptation of language models in the financial sector. Future research directions include optimizing textual modality parameters, such as the time window, sentiment, and chronological order of news messages.

多模态股市预测新闻分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。