arXiv:2605.30652cs.LG2026-05

用密集嵌入替代情感分值,提升股市预测准确性

Bridging the Gap Between Natural Language and Market Dynamics via High-Dimensional Representation Learning

论文配图:Bridging the Gap Between Natural Language and Market Dynamics via High-Dimensional Representation Learning
图 1 · 摘自论文原文
  • 用FinBERT生成高维嵌入,替代传统情感评分
  • 自定义孪生网络优化嵌入,显著提升预测效果
  • 适合关注金融文本建模与量化交易的研究者

传统多模态金融预测依赖离散的情感分数,难以捕捉财经新闻的细微语义。本文提出用密集的FinBERT嵌入替代离散极性评分,融入基于Transformer的预测架构。在FNSPID数据集上对比了原始嵌入、注意力加权聚合及自定义孪生网络等多种策略。尽管注意力机制因金融数据信噪比低而表现不佳,但孪生网络优化的嵌入显著优于标量基线和原始嵌入方法,证明保留高维叙事上下文能有效提升短期股价变动的预测精度。

原文摘要 · Abstract (English)

Traditional multi-modal financial forecasting often relies on scalar sentiment scores, which fail to capture the nuances of financial news. To address this information loss, this paper explores high-dimensional representation learning by replacing discrete polarity ratings with dense FinBERT embeddings within a Transformer-based forecasting architecture. We benchmarked various embedding strategies on the FNSPID dataset, including raw embeddings, attention-weighted aggregation, and a custom Siamese network. While the attention-based mechanism struggled with the low signal-to-noise ratio typical of financial data, the integration of Siamese-optimized embeddings outperformed both the scalar baseline and raw embedding approaches, demonstrating that preserving high-dimensional narrative context yields improved predictive accuracy for short-term stock price movements.

金融预测文本嵌入TransformerFinBERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。