融合新闻与股价的多模态模型,提升金融预测准确率
Multimodal Language Models with Modality-Specific Experts for Financial Forecasting from Interleaved Sequences of Text and Time Series
- 用特定专家模块分别处理文本和时间序列数据
- 在多个基准上达到领先性能,经济收益显著
- 可解释性分析揭示关键信息对齐机制
文本与时间序列数据为金融市场提供互补视角:新闻文章提供公司事件的叙述背景,而股价则反映市场对这些事件的反应。尽管二者互补,但有效整合交错的多模态数据以提升预测性能仍具挑战。本文提出一种统一神经架构,通过模态特定专家建模交错序列,使模型能学习独特的时序模式,同时实现跨模态联合推理,并保留预训练语言模型的理解能力。为进一步增强多模态理解,引入基于显著标记加权的跨模态对齐框架,聚焦最具信息量的标记进行表示对齐。在大规模金融预测任务上验证了该方法的有效性,在多种强基线(包括单模态与多模态)中取得最优表现。我们开发了一种可解释性方法,揭示时序-上下文价值关系,支持跨模态对齐目标的设计。最终,这些改进在投资模拟中转化为可观的经济收益。
原文摘要 · Abstract (English)
Text and time series data offer complementary views of financial markets: news articles provide narrative context about company events, while stock prices reflect how markets react to those events. However, despite their complementary nature, effectively integrating these interleaved modalities for improved forecasting remains challenging. In this work, we propose a unified neural architecture that models these interleaved sequences using modality-specific experts, allowing the model to learn unique time series patterns, while still enabling joint reasoning across modalities and preserving pretrained language understanding capabilities. To further improve multimodal understanding, we introduce a cross-modal alignment framework with a salient token weighting mechanism that learns to align representations across modalities with a focus on the most informative tokens. We demonstrate the effectiveness of our approach on a large-scale financial forecasting task, achieving state-of-the-art performance across a wide variety of strong unimodal and multimodal baselines. We develop an interpretability method that reveals insights into the value of time series-context and reinforces the design of our cross-modal alignment objective. Finally, we demonstrate that these improvements translate to meaningful economic gains in investment simulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。