arXiv:2510.10744stat.MLcs.IT2025-10NeurIPS

用信息论量化序列数据可学习性,揭示模型性能上限。

How Patterns Dictate Learnability in Sequential Data

  • 以过去与未来间的互信息定义预测信息,构建信息论学习曲线。
  • 证明时间模式缺失时,再优模型也无法突破数据本身的信息极限。
  • 适用于评估模型是否合适、分析数据复杂度,适合做序列建模前的预判。

序列数据(从金融时间序列到自然语言)推动了自回归模型的广泛应用。然而,这些算法依赖数据中隐藏的模式,而模式识别常需人工经验。误判模式会导致模型误设,增加泛化误差并降低性能。近期提出的演化模式(EvoRate)指标通过下一数据点与其历史之间的互信息,指导回归阶数估计和特征选择。基于此思想,本文提出一个基于预测信息的一般框架:即过去与未来之间的互信息 $I(X_{past}; X_{future})$。该量自然定义了一个信息论学习曲线,量化了随着观察窗口增长,可用的预测信息量。我们证明,时间模式的存在与否从根本上限制了序列模型的可学习性:即使最优预测器也无法超越数据内在的信息极限。通过合成数据实验验证了该框架的有效性,其能评估模型适配性、量化数据固有复杂度,并揭示序列数据中可解释的结构。

原文摘要 · Abstract (English)

Sequential data - ranging from financial time series to natural language - has driven the growing adoption of autoregressive models. However, these algorithms rely on the presence of underlying patterns in the data, and their identification often depends heavily on human expertise. Misinterpreting these patterns can lead to model misspecification, resulting in increased generalization error and degraded performance. The recently proposed evolving pattern (EvoRate) metric addresses this by using the mutual information between the next data point and its past to guide regression order estimation and feature selection. Building on this idea, we introduce a general framework based on predictive information, defined as the mutual information between the past and the future, $I(X_{past}; X_{future})$. This quantity naturally defines an information-theoretic learning curve, which quantifies the amount of predictive information available as the observation window grows. Using this formalism, we show that the presence or absence of temporal patterns fundamentally constrains the learnability of sequential models: even an optimal predictor cannot outperform the intrinsic information limit imposed by the data. We validate our framework through experiments on synthetic data, demonstrating its ability to assess model adequacy, quantify the inherent complexity of a dataset, and reveal interpretable structure in sequential data.

信息论序列建模可学习性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。