arXiv:2503.10957cs.LGcs.AI2025-03被引 2

用Twitter预训练的BERTweet模型预测股票走势,效果优于现有方法。

Predicting Stock Movement with BERTweet and Transformers

  • 采用BERTweet和Transformer架构处理社交媒体文本数据
  • 在Stocknet数据集上达到新的马修斯相关系数基准
  • 无需额外数据源,适合金融自然语言处理研究者

将深度学习与计算智能应用于金融领域是当前学术界和工业界的热门方向,但高波动性和非平稳性给机器学习模型带来挑战,尤其是对参数量大的深度模型。近期研究通过结合社交媒体文本的自然语言处理技术,提升仅依赖历史价格数据的模型表现,受到广泛关注。此前工作已通过双向GRU、变分自编码器、词与文档嵌入、自注意力、图注意力及对抗训练等技术实现领先性能。本文证明了专门针对推特语料预训练的BERTweet及其变换器架构的有效性,在不使用辅助数据源的情况下,于Stocknet数据集上取得了具有竞争力的表现,并建立了新的马修斯相关系数基准。

原文摘要 · Abstract (English)

Applying deep learning and computational intelligence to finance has been a popular area of applied research, both within academia and industry, and continues to attract active attention. The inherently high volatility and non-stationary of the data pose substantial challenges to machine learning models, especially so for today's expressive and highly-parameterized deep learning models. Recent work has combined natural language processing on data from social media to augment models based purely on historic price data to improve performance has received particular attention. Previous work has achieved state-of-the-art performance on this task by combining techniques such as bidirectional GRUs, variational autoencoders, word and document embeddings, self-attention, graph attention, and adversarial training. In this paper, we demonstrated the efficacy of BERTweet, a variant of BERT pre-trained specifically on a Twitter corpus, and the transformer architecture by achieving competitive performance with the existing literature and setting a new baseline for Matthews Correlation Coefficient on the Stocknet dataset without auxiliary data sources.

金融AIBERTweet文本预测Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。