arXiv:2603.12451cs.LG2026-03

用合成数据提升上下文质量,让模型真正用好领域知识进行预测。

Overcoming the Modality Gap in Context-Aided Forecasting

  • 通过半合成方法生成可验证的高质量上下文,改善数据缺陷
  • 构建700万条带上下文的时间序列数据集,验证效果显著优于旧模型
  • 证明预测性能瓶颈在数据而非模型结构,适合做智能预测研究者参考

上下文辅助预测(CAF)有望融合领域知识与前瞻性信息,使AI系统超越传统统计方法。然而近期实证研究发现,多模态模型常不及单模态模型表现。我们推测此现象源于现有数据集中上下文质量差,且难以验证。为此,我们提出一种半合成数据增强方法,生成既描述时序动态又可验证互补数值历史的上下文。该方法支持大规模数据集构建,形成包含700万条上下文增强时间序列窗口的CAF-7M数据集,并配有严格验证的测试集。实验表明,半合成预训练能有效迁移到真实场景,且明确证明了上下文被有效利用。结果表明,数据质量而非架构限制才是制约上下文辅助预测的主要瓶颈。

原文摘要 · Abstract (English)

Context-aided forecasting (CAF) holds promise for integrating domain knowledge and forward-looking information, enabling AI systems to surpass traditional statistical methods. However, recent empirical studies reveal a puzzling gap: multimodal models often fail to outperform their unimodal counterparts. We hypothesize that this underperformance stems from poor context quality in existing datasets, as verification is challenging. To address these limitations, we introduce a semi-synthetic data augmentation method that generates contexts both descriptive of temporal dynamics and verifiably complementary to numerical histories. This approach enables massive-scale dataset creation, resulting in CAF-7M, a corpus of 7 million context-augmented time series windows, including a rigorously verified test set. We demonstrate that semi-synthetic pre-training transfers effectively to real-world evaluation, and show clear evidence of context utilization. Our results suggest that dataset quality, rather than architectural limitations, has been the primary bottleneck in context-aided forecasting.

时间序列预测上下文增强数据增强多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。