辅助上下文能否提升时间序列预测,取决于两个关键条件是否满足。
When Does Context Routing Help? A Systematic Study of Multi-Modal Fusion in Time Series Forecasting

- 提出判断上下文是否有效的两个数据层面标准:低自相关性和上下文含目标信息。
- 当双条件满足时,文本调制专家模型使均方误差降低44%;否则效果归零。
- 适用于评估多模态融合有效性,尤其适合研究者和工业界验证上下文价值。
多模态时间序列预测方法通过复杂的融合机制整合辅助上下文以提升预测性能。尽管大量工作报告显著提升,但难以区分是真正利用了上下文,还是架构带来的偶然效应。本文聚焦一个可检验的问题:辅助上下文何时能帮助预测?我们识别出两个必须同时满足的数据集级条件:(1) 目标变量不被最后值捷径主导(自相关系数rho_h较低);(2) 上下文携带历史之外的目标信息(非零条件互信息delta;当delta=0时,任何预测器都无法受益——分布无关结论)。在包含14.3B参数的MoME模型(6个数据集,10次随机种子)及单骨干测试平台中实现四种融合机制(5个数据集)的受控实验表明,当双条件成立时,文本调制专家模块带来显著的均方误差下降;任一条件失效,则贡献降为调制路径容量下限,无上下文特异性信号。通过干预实验证实因果性:在MoME中添加捷径使路由贡献下降77%-93%;逐步破坏上下文质量使上下文收益从+44%转为负值。在27个Monash Archive数据集上验证了自相关诊断的有效性。我们提供了一个校准的预训练诊断工具,在测试数据集中未出现假阳性。我们明确指出证据的不对称性:负面结果广泛可靠,而较大正向效应仅来自单一模型族(MoME),方向一致性由测试平台支持。
原文摘要 · Abstract (English)
Multi-modal time series forecasting methods integrate auxiliary context into temporal predictions through increasingly sophisticated fusion mechanisms. A growing body of work reports substantial gains, yet it is often unclear whether they reflect genuine use of the context or incidental architectural effects. We ask a narrower, checkable question: when can auxiliary context help a forecaster at all? We identify two dataset-level conditions that must both hold: (1) the target is not dominated by a last-value shortcut (low autocorrelation rho_h), and (2) the context carries information about the target beyond history (non-zero conditional mutual information delta; when delta=0 no predictor can benefit---a distribution-free result). Through controlled experiments on MoME (a 14.3B-parameter mixture-of-experts model, 6 datasets, 10 seeds) and four additional fusion mechanisms implemented within a single-backbone testbed (5 datasets), we find that when both conditions hold, text-conditioned expert modulation contributes a sizeable MSE reduction; when either fails, the contribution collapses to the capacity floor of the modulation pathway and carries no context-attributable signal. We establish causality through two interventions: adding a shortcut to MoME suppresses routing contribution by 77-93% across 3 datasets; progressively corrupting context quality drives the context-specific benefit from +44% to negative. We validate the autocorrelation component of our diagnostic on 27 Monash Archive datasets. We provide a calibrated pre-training diagnostic that, on the datasets we test, yields no false positives in well-powered settings. We are explicit about the asymmetry of our evidence: the negative arm is broadly reliable, while the large positive magnitudes come from a single model family (MoME) and are corroborated only in direction by the testbed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。