Toto模型用1.5亿参数提升时序预测,专为可观测性数据优化。
This Time is Different: An Observability Perspective on Time Series Foundation Models
- 采用解码器架构+创新设计,适配多变量可观测时序数据挑战
- 在350万观测点的BOOM基准上达到当前最优性能
- 适用于需要高精度时序预测的运维与监控场景
我们提出Toto,一个拥有1510万参数的时序预测基础模型。Toto采用现代解码器架构,并引入针对多变量可观测时序数据特性的架构改进。其预训练语料库包含可观测性数据、公开数据集和合成数据,规模为领先时序基础模型的4-10倍。此外,我们构建了BOOM,一个涵盖2,807个真实世界时序、共3.5亿条观测的大规模基准。Toto与BOOM的数据均来自Datadog内部遥测与可观测性指标。大量实验表明,Toto在BOOM及主流通用时序预测基准上均达到顶尖表现。Toto的模型权重、推理代码与评估脚本,以及BOOM的数据与评估代码,均已开源,许可协议为Apache 2.0,可通过Hugging Face与GitHub获取。
原文摘要 · Abstract (English)
We introduce Toto, a time series forecasting foundation model with 151 million parameters. Toto uses a modern decoder-only architecture coupled with architectural innovations designed to account for specific challenges found in multivariate observability time series data. Toto's pre-training corpus is a mixture of observability data, open datasets, and synthetic data, and is 4-10$\times$ larger than those of leading time series foundation models. Additionally, we introduce BOOM, a large-scale benchmark consisting of 350 million observations across 2,807 real-world time series. For both Toto and BOOM, we source observability data exclusively from Datadog's own telemetry and internal observability metrics. Extensive evaluations demonstrate that Toto achieves state-of-the-art performance on both BOOM and on established general purpose time series forecasting benchmarks. Toto's model weights, inference code, and evaluation scripts, as well as BOOM's data and evaluation code, are all available as open source under the Apache 2.0 License available at https://huggingface.co/Datadog/Toto-Open-Base-1.0 and https://github.com/DataDog/toto.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。