提出可解释的多时序商品期货预测模型,解决训练与推理不一致问题。
Sparse Latent Factor Forecaster (SLFF) with Iterative Inference for Transparent Multi-Horizon Commodity Futures Prediction
- 通过稀疏编码与迭代优化提升潜在因子质量
- 1-5天预测误差显著低于神经网络基线
- 因子稳定且与经济基本面相关,适合金融从业者参考
隐变量预测模型中的变分推断存在部署差距:测试时编码器仅能近似训练时经过目标优化的潜在表示,却无法访问未来目标。这引入额外预测误差并影响可解释性。本文提出稀疏潜在因子预测器(SLFF),通过三方面改进:(i) 基于L1正则的稀疏编码目标,生成低维潜在因子;(ii) 在训练中采用类似LISTA的展开式近端梯度下降进行迭代优化;(iii) 编码器对齐,确保摊销输出与优化解一致。在解码器线性假设下,我们推导出基于编码器-优化器距离的摊销差距上界,并证明在弱条件下具有收敛率;实证检验显示该上界对部署的MLP解码器具有预测能力。为防止高频数据泄露,引入基于发布日历和历史宏观经济数据的信息集感知协议。可解释性通过三阶段框架形式化:稳定性(跨种子的Procrustes对齐)、驱动因素有效性(剔除回归与可观测变量关联)、行为一致性(反事实分析与事件研究)。以铜、WTI原油、黄金期货(2005–2025)为测试基准,SLFF在1天和5天预测上显著优于神经网络基线,生成的稀疏因子在不同种子间稳定,且与可观测经济基本面相关(可解释性为相关性,非因果)。代码、数据集清单、诊断结果与实验资产均已公开。
原文摘要 · Abstract (English)
Amortized variational inference in latent-variable forecasters creates a deployment gap: the test-time encoder approximates a training-time optimization-refined latent, but without access to future targets. This gap introduces unnecessary forecast error and interpretability challenges. In this work, we propose the Sparse Latent Factor Forecaster with Iterative Inference (SLFF), addressing this through (i) a sparse coding objective with L1 regularization for low-dimensional latents, (ii) unrolled proximal gradient descent (LISTA-style) for iterative refinement during training, and (iii) encoder alignment to ensure amortized outputs match optimization-refined solutions. Under a linearized decoder assumption, we derive a design-motivating bound on the amortization gap based on encoder-optimizer distance, with convergence rates under mild conditions; empirical checks confirm the bound is predictive for the deployed MLP decoder. To prevent mixed-frequency data leakage, we introduce an information-set-aware protocol using release calendars and vintage macroeconomic data. Interpretability is formalized via a three-stage protocol: stability (Procrustes alignment across seeds), driver validity (held-out regressions against observables), and behavioral consistency (counterfactuals and event studies). Using commodity futures (Copper, WTI, Gold; 2005--2025) as a testbed, SLFF demonstrates significant improvements over neural baselines at 1- and 5-day horizons, yielding sparse factors that are stable across seeds and correlated with observable economic fundamentals (interpretability remains correlational, not causal). Code, manifests, diagnostics, and artifacts are released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。