首个动态时间序列基准,测试模型在真实世界中的长期表现。
LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models

- 用实时数据流持续评估模型,而非固定历史窗口
- 11个领域17个数据集显示静态排名大幅变化
- 适合关注模型长期鲁棒性的研究者
时间序列基础模型(TSFMs)近期成为跨领域零样本预测的有前途范式。然而,现有评估协议主要依赖固定历史测试窗口的静态基准。这类基准虽提供有价值的基础快照,但仅评估模型在固定历史上的平均性能,无法捕捉其在随季节变化、分布漂移和突发事件持续演化的现实环境中表现。为此,我们提出LiveHouse-TS,首个面向时间序列基础模型的开放世界动态基准框架。通过在开放世界环境中对真实未来数据进行预序评估,LiveHouse-TS将时间序列基准从快照准确率转向持续的时间有效性。该框架不只是一次性排行榜,而是持续运行的时序基础设施,用于探索关键的长期科学问题:模型排名能否长期维持?哪些模型在分布漂移下仍具真正鲁棒性?在11个领域17个数据集上的广泛流式评估表明,静态排名在动态协议下发生剧烈重排。
原文摘要 · Abstract (English)
Time Series Foundation Models (TSFMs) have recently emerged as a highly promising paradigm for cross-domain zero-shot forecasting. However, existing evaluation protocols predominantly rely on static benchmarks with fixed historical test windows. While these benchmarks provide a valuable baseline snapshot, they evaluate an average performance on a fixed history, failing to capture how models behave in continuously evolving real-world environments characterized by seasonal variations, distribution shifts, and unexpected events. To bridge this gap, we introduce LiveHouse-TS, the first open-world living benchmark infrastructure for TSFMs. By evaluating models prequentially on real future data in open-world environments, LiveHouse-TS shifts time series benchmarking from snapshot accuracy to continuous temporal validity. Rather than acting as a one-off leaderboard, our infrastructure serves as a continuous time series infrastructure designed to explore vital, long-term scientific questions: Can model rankings be maintained over the long term? Which models remain genuinely robust under distribution shifts? Extensive streaming evaluations across 11 domains with 17 datasets demonstrate that static rankings undergo a dramatic reshuffling under a live protocol.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。