合成数据组合比选单一生成器更重要,能显著提升时间序列模型预训练效果。
Mix, Don't Pick: Why Synthetic Corpus Composition Matters for Time Series Foundation Model Pretraining
- 用等权重混合所有生成器构建合成语料库
- 混合语料使预测误差降低至单个最佳生成器的水平
- 适合关注时间序列预训练策略的研究者
选择错误的合成生成器会严重影响时间序列基础模型的预训练效果:在相同训练预算下,表现最好与最差的生成器之间预测误差相差高达2倍。然而,当前领域缺乏系统性的生成器选择方法。更严重的是,生成器的有效性随模型架构变化而波动——在Chronos-T5-Mini和Moirai-Small两个模型上评估11类生成器时发现,哪些生成器有效取决于具体架构。与其解决生成器选择难题,不如绕开它:将所有生成器等权混合,其性能可匹配甚至超越单一最优生成器;再将该混合语料与真实数据结合,整体预训练效果最强。因此,合成预训练本质上是语料组合问题,而非生成器选择问题,组合策略应针对不同模型族单独验证,不能盲目迁移。
原文摘要 · Abstract (English)
Choosing the wrong synthetic generator for time-series foundation model pretraining is costly: under identical training budgets, the best and worst generators produce up to a $2\times$ gap in forecasting error, yet the field has no principled way to make this choice. The problem is compounded by the fact that generator rankings are not stable across architectures: across 11 generator families evaluated on Chronos-T5-Mini and Moirai-Small trained from scratch, we find that which generators are useful depends on the model architecture. Rather than solving the generator selection problem, we sidestep it: a simple equal-weight mixture of all generators matches or beats the best individual generator for both architectures, and composing this mixture with real data yields the strongest pretraining corpora overall. Synthetic pretraining is therefore a corpus composition problem, not a generator selection problem, and composition choices should be validated per model family rather than assumed to transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。