arXiv:2509.21002cs.LGcs.AI2025-09

用无损压缩评估时间序列模型,揭示传统任务遗漏的分布缺陷

Lossless Compression: A New Benchmark for Time Series Model Evaluation

  • 以香农编码定理为基础,将最优压缩长度等价为负对数似然
  • 在多个数据集上发现主流模型在分布建模上存在被传统任务忽略的弱点
  • 开源框架TSCom-Bench可快速测试模型压缩性能,适合评估生成能力

时间序列模型的评估长期聚焦于预测、填补缺失、异常检测和分类四类任务。这些任务虽推动了进展,但主要衡量特定任务表现,无法严格检验模型是否捕捉数据的完整生成分布。本文提出以无损压缩作为新评估范式,基于香农信源编码定理,建立最优压缩长度与负对数似然之间的直接等价关系,提供统一且严格的信道信息论标准。我们定义标准化评估协议与指标,并开源全面的评估框架TSCom-Bench,支持将时间序列模型快速作为无损压缩的骨干。在多种数据集上对TimeXer、iTransformer、PatchTST等先进模型的实验表明,压缩方法能揭示经典基准未发现的分布建模缺陷。研究结果表明,无损压缩是一种有原则的新任务,可补充并扩展现有时间序列建模评估体系。

原文摘要 · Abstract (English)

The evaluation of time series models has traditionally focused on four canonical tasks: forecasting, imputation, anomaly detection, and classification. While these tasks have driven significant progress, they primarily assess task-specific performance and do not rigorously measure whether a model captures the full generative distribution of the data. We introduce lossless compression as a new paradigm for evaluating time series models, grounded in Shannon's source coding theorem. This perspective establishes a direct equivalence between optimal compression length and the negative log-likelihood, providing a strict and unified information-theoretic criterion for modeling capacity. Then We define a standardized evaluation protocol and metrics. We further propose and open-source a comprehensive evaluation framework TSCom-Bench, which enables the rapid adaptation of time series models as backbones for lossless compression. Experiments across diverse datasets on state-of-the-art models, including TimeXer, iTransformer, and PatchTST, demonstrate that compression reveals distributional weaknesses overlooked by classic benchmarks. These findings position lossless compression as a principled task that complements and extends existing evaluation for time series modeling.

时间序列无损压缩模型评估信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。