用频谱图和时序数据融合预测商品价格,提升准确性与鲁棒性。
Multimodal Forecasting for Commodity Prices Using Spectrogram-Based and Time Series Representations
- 将价格序列转为梅尔小波频谱图,结合视觉变压器提取频率特征
- 融合宏观经济指标的时序编码,跨模态注意力建模变量间关系
- 在多个商品价格预测任务中超越7个基线模型,适合金融时间序列研究者
多变量时间序列预测因复杂的变量间依赖关系及异质外部影响而具挑战性。本文提出频谱增强多模态融合(SEMF)方法,结合频谱与时序表示以实现更准确、鲁棒的预测。目标时间序列经梅尔小波变换生成频谱图,由视觉变压器编码器提取局部化、频率感知特征;同时,金融指标与宏观经济信号通过Transformer编码以捕捉时序依赖与多变量动态。双向交叉注意力模块将双模态信息融合为统一表征,保留各自信号特性并建模跨模态关联。在多个商品价格预测任务中,SEMF在多个预测时长与评估指标上持续优于7个竞争基线模型。结果表明,多模态融合与基于频谱的编码能有效捕捉复杂金融时间序列中的多尺度模式。
原文摘要 · Abstract (English)
Forecasting multivariate time series remains challenging due to complex cross-variable dependencies and the presence of heterogeneous external influences. This paper presents Spectrogram-Enhanced Multimodal Fusion (SEMF), which combines spectral and temporal representations for more accurate and robust forecasting. The target time series is transformed into Morlet wavelet spectrograms, from which a Vision Transformer encoder extracts localized, frequency-aware features. In parallel, exogenous variables, such as financial indicators and macroeconomic signals, are encoded via a Transformer to capture temporal dependencies and multivariate dynamics. A bidirectional cross-attention module integrates these modalities into a unified representation that preserves distinct signal characteristics while modeling cross-modal correlations. Applied to multiple commodity price forecasting tasks, SEMF achieves consistent improvements over seven competitive baselines across multiple forecasting horizons and evaluation metrics. These results demonstrate the effectiveness of multimodal fusion and spectrogram-based encoding in capturing multi-scale patterns within complex financial time series.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。