金融收益预测中,输出头比主干网络更重要,混合分布头能更好捕捉极端风险。
Heads, Not Backbones: Output Heads Dominate Architectures on Fat-Tailed Returns

- 用三种输出头对比四种主干网络,发现头的选择影响远大于主干差异。
- 混合高斯头相比单高斯头,CRPS降低2.4%,危机时期提升达13.9%。
- 短时预测下头主导性能,长时预测则主干重新占优,适合风险敏感场景。
在短期肥尾金融收益的深度预测流程中,主干架构与输出头哪个更重要?我们比较了四种现代主干(TimesNet、DLinear、N-BEATS、iTransformer)搭配三种输出头:点预测头、单高斯密度头、含K=4分量的高斯混合密度头。在标普500月度对数收益(1871-2023)上,采用锚定滚动验证,三类头形成严格梯度:从点预测到单高斯,CRPS改善约1.3%;再到混合高斯,再降2.4%。主干间切换对点头影响不足1.5%,对主干均值轴影响也小;虽密度头下主干差异最大达5.1%(由N-BEATS驱动),但头之间的差距(3.7个百分点)仍占主导。基于平方误差的模型置信集未在5%水平排除任一12种组合,说明头仅在分布指标(CRPS、pinball、覆盖率)上区分模型,而非平方误差。混合头在高波动期(如1970年代滞胀期,h=12时)相对于单高斯提升达13.9%,证实其捕捉尾部风险的能力超越单峰高斯。该结果具有时序依赖性:短时预测中头主导,长时预测(h≥6)主干重获优势,此分裂现象在经典基线中亦可验证。
原文摘要 · Abstract (English)
In a deep forecasting pipeline for fat-tailed financial returns at short horizons, which matters more - the backbone architecture or the output head? We compare four modern backbones (TimesNet, DLinear, N-BEATS, iTransformer) under three output heads: a point head, a single-Gaussian density head, and a Gaussian mixture density head with K=4 components. On S and P 500 monthly log-returns (1871-2023) under anchored walk-forward validation, the three heads form a strict gradient: switching from point to Gaussian improves CRPS by about 1.3 percent; switching from Gaussian to mixture adds a further about 2.4 percent. Switching between backbones, in contrast, changes CRPS by less than 1.5 percent on the point-head row and on the backbone-mean axis; density-head backbone spread is larger (up to 5.1 percent on the h=1 Gaussian row, driven by N-BEATS) but the head gradient (3.7 percentage points) still dominates. The Model Confidence Set on squared errors does not exclude any of the 12 variants at the 5 percent level: the head separates them only on distributional metrics (CRPS, pinball, coverage), not on squared error. The mixture head incremental value over a single Gaussian is largest in the highest-volatility regimes (13.9 percent in 1970s stagflation at h=12), confirming the mixture captures tail risk beyond what a unimodal Gaussian can express. The picture is horizon-dependent: the head dominates at short horizons, but at long horizons (h >= 6) the backbone re-takes the lead - an h-split we document against classical baselines (section 5.1). We conclude that on fat-tailed returns at short horizons, the head dominates the backbone, and the mixture distribution adds genuine value over a single Gaussian during crisis periods when risk-management decisions actually matter.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。